Search bioRxiv⌕ Search

Biology subjects

Weaver, T. D.

Publications and source records attributed to Weaver, T. D..

3 recordsLinked to original sources

Improving GWAS performance in underrepresented groups by appropriate modeling of genetics, environment, and sociocultural factors

Genome-wide association studies (GWAS) and polygenic score (PGS) development are typically constrained by the data available in biobank repositories in which European cohorts are vastly overrepresented. Here, we increase the utility of non-European participant data within the UK Biobank (UKB) by characterizing the genetic affinities of UKB participants who self-identify as Bangladeshi, Indian, Pakistani, "White and Asian" (WA), and "Any Other Asian" (AOA), towards creating a more robust South Asian sample size for future genetic analyses. We assess the relationships between genetic structure and self-selected ethnic identities and use consistent patterns of clustering in the dataset to train a support vector machine (SVM). The SVM was utilized to reassign n = 1,853 AOA and WA participants at the subcontinental level, and increase the sample size of the UKB South Asian group by 1,381 additional participants. We further leverage these samples to assess GWAS performance and PGS development. We include environmental covariates in the height GWAS by implementing a rigorous covariate selection procedure, and compare the outputs of two GWAS models: GWASnull and GWASenv. We show that PGS performance derived from both GWAS models yield comparable prediction to PGS models developed with an order of magnitude larger training, and environmentally-adjusted PGS models reduce the sex-bias in predictive performance. In summary, we demonstrate how GWAS performance can be improved by leveraging ambiguous ethnicity codes, ancestry matched imputation panels, and including environmental covariates.

genetics↗

Developing an automated skeletal phenotyping pipeline to leverage biobank-level medical imaging databases

ObjectivesCollecting skeletal measurements from medical imaging databases remains a tedious task, limiting the research utility of biobank-level data. Here we present an automated phenotyping pipeline for obtaining skeletal measurements from DXA scans and compare its performance to manually collected measurements. Materials and MethodsA pipeline that extends and modifies the Advanced Normalization Tools (ANTs) framework was developed on 341 whole-body DXA scans of UK Biobank South Asian participants. A set of 10 measurements throughout the skeleton was automatically obtained via this process, and the performance of the method was tested on 20 additional DXA images by calculating percent error and concordance correlation coefficients (CCC) for manual and automated measurements. Stature was then regressed on the automated femoral and tibia lengths and compared to published stature regressions to further assess the reliability of the automated measurements. ResultsBased on percent error and CCC, the performance of the automated measurements falls into three categories: poor (sacral and acetabular breadths), variable (trunk length, upper thoracic breadth, and innominate height), and high (maximum pelvic aperture breadth, bi-iliac breadth, femoral maximum length, and tibia length). Stature regression plots indicate that the automated measurements reflect realistic body proportions and appear consistent with published data reflecting these relationships in South Asian populations. DiscussionBased on the performance of this pipeline, a subset of measurements can be reliably extracted from DXA scans, greatly expanding the utility of biobank-level data for biological anthropologists and medical researchers.

bioinformatics↗

A weakly structured stem for human origins in Africa

While it is now broadly accepted that Homo sapiens originated within Africa, considerable uncertainty surrounds specific models of divergence and migration across the continent. Progress is hampered by a paucity of fossil and genomic data, as well as variability in prior divergence time estimates. Here we use linkage disequilibrium and diversity-based statistics, optimized for rapid, complex demographic inference to discriminate among such models. We infer detailed demographic models for populations across Africa, including representatives from eastern and western groups, as well as 44 newly whole-genome sequenced individuals from the Nama (Khoe-San). Despite the complexity of African population history, contemporary population structure dates back to Marine Isotope Stage (MIS) 5. The earliest population divergence among contemporary populations occurs 120-135ka, between the Khoe-San and other groups. Prior to the divergence of contemporary African groups, we infer long-lasting structure between two or more weakly differentiated ancestral Homo populations connected by gene flow over hundreds of thousands of years (i.e. a weakly structured stem). We find that weakly structured stem models provide more likely explanations of polymorphism that had previously been attributed to contributions from archaic hominins in Africa. In contrast to models with archaic introgression, we predict that fossil remains from coexisting ancestral populations should be morphologically similar. Despite genetic similarity between these populations, an inferred 1-4% of genetic differentiation among contemporary human populations can be attributed to genetic drift between stem populations. We show that model misspecification explains variation in previous divergence time estimates and argue that studying a suite of models is key to robust inferences about deep history.

genomics↗