Search bioRxiv⌕ Search

Biology subjects

Baya, N.

Publications and source records attributed to Baya, N..

4 recordsLinked to original sources

On the Genes, Genealogies, and Geographies of Quebec

Population genetic models only provide coarse representations of real-world ancestry. We use a pedigree compiled from four million parish records and genotype data from 2,276 French and 20,451 French Canadian (FC) individuals, to finely model and trace FC ancestry through space and time. The loss of ancestral French population structure and the appearance of spatial and regional structure highlights a wide range of population expansion models. Geographic features shaped migrations throughout, and we find enrichments for migration, genetic and genealogical relatedness patterns within river networks across Quebec regions. Finally, we provide a freely accessible simulated whole-genome sequence dataset with spatiotemporal metadata for 1,426,749 individuals reflecting intricate FC population structure. Such realistic populations-scale simulations provide new opportunities to investigate population genetics at an unprecedented resolution. Lay SummaryWe all share common ancestors ranging from a couple generations ago to hundreds of thousands of years ago. The genetic differences between individuals today mostly depends on how closely related they are. The only problem is that the actual genealogies that relate all of us are often forgotten over time. Some geneticists have tried to come up with simple models of our shared ancestry but they dont really explain the full, rich history of humanity. Our study uses a multi-institutional project in Quebec that has digitized parish records into a single unified genealogical database that dates back to the arrival of the first French settlers four hundred years ago. This genealogy traces the ancestry of millions of French-Canadian and we have used it to build a very high resolution genetic map. We used this genetic map to study in detail how certain historical events, and landscapes have influenced the genomes of French-Canadians today. One-Sentence SummaryWe present an accurate and high resolution spatiotemporal model of genetic variation in a founder population.

genomics↗

Patterns of item nonresponse behavior to survey questionnaires are systematic and have a genetic basis

Response to survey questionnaires is vital for social and behavioral research, and most analyses assume full and accurate response by survey participants. However, nonresponse is common and impedes proper interpretation and generalizability of results. We examined item nonresponse behavior across 109 questionnaire items from the UK Biobank (UKB) (N=360,628). Phenotypic factor scores for two participant-selected nonresponse answers, "Prefer not to answer" (PNA) and "I dont know" (IDK), each predicted participant nonresponse in follow-up surveys, controlling for education and self-reported general health. We performed genome-wide association studies on these factors and identified 39 genome-wide significant loci, and further validated these effects with polygenic scores in an independent study (N=3,414), gaining information that we could not have had from phenotypic data alone. PNA and IDK were highly genetically correlated with one another and with education, health, and income, although unique genetic effects were also observed for both PNA and IDK. We discuss how these effects may bias studies of traits correlated with nonresponse and how genetic analyses can further enhance our understanding of nonresponse behaviors in survey research, for instance by helping to correct for nonresponse bias.

genetics↗

Analysis of genetic dominance in the UK Biobank

Classical statistical genetic theory defines dominance as a deviation from a purely additive effect. Dominance is well documented in model organisms and plant/animal breeding; outside of rare monogenic traits, however, evidence in humans is limited. We evaluated dominance effects in >1,000 phenotypes in the UK Biobank through GWAS, identifying 175 genome-wide significant loci (P < 4.7 x 10-11). Power to detect non-additive loci is low: we estimate a 20-30 fold increase in sample size is required to detect dominance loci to significance levels observed at additive loci. By deriving a new dominance form of LD-score regression, we found no evidence of a dominance contribution to phenotypic variance tagged by common variation genome-wide (median fraction 5.73 x 10-4). We introduce dominance fine-mapping to explore whether the more rapid decay of dominance linkage disequilibrium can be leveraged to find causal variants. These results provide the most comprehensive assessment of dominance trait variation in humans to date.

genetics↗

Genetic analyses identify widespread sex-differential participation bias

Genetic association results are often interpreted with the assumption that study participation does not affect downstream analyses. Understanding the genetic basis of this participation bias is challenging as it requires the genotypes of unseen individuals. However, we demonstrate that it is possible to estimate comparative biases by performing GWAS contrasting one subgroup versus another. For example, we show that sex exhibits autosomal heritability in the presence of sex-differential participation bias. By performing a GWAS of sex in ~3.3 million males and females, we identify over 158 autosomal loci significantly associated with sex and highlight complex traits underpinning differences in study participation between sexes. For example, the body mass index (BMI) increasing allele at the FTO locus was observed at higher frequency in males compared to females (OR 1.02 [1.02-1.03], P=4.4x10-36). Finally, we demonstrate how these biases can potentially lead to incorrect inferences in downstream analyses and propose a conceptual framework for addressing such biases. Our findings highlight a new challenge that genetic studies may face as sample sizes continue to grow.

genetics↗