Search bioRxiv⌕ Search

Biology subjects

Honorato-Mauer, J.

Publications and source records attributed to Honorato-Mauer, J..

3 recordsLinked to original sources

Tractor Workflow Pipeline: A Scalable Nextflow Framework for Local Ancestry-Aware Genome-Wide Association Studies

The routine exclusion of admixed individuals from traditional Genome-Wide Association Studies (GWAS) due to concerns about spurious associations has hindered genetic analyses involving multiple ancestries. Tractor GWAS addresses this issue by incorporating local ancestry into its analysis, empowering identification of ancestry-enriched hits and generating ancestry-specific summary statistics. However, Tractor requires accurate genomic phasing and local ancestry inference as prerequisite steps, which requires additional bioinformatics expertise and decision points regarding reference panel setup. To streamline, harmonize, and automate this process, we present a scalable Nextflow workflow that integrates all necessary steps, minimizing the need for manual intervention while remaining modular and customizable. The workflow supports multiple commonly used tools and offers flexibility in how Tractor is implemented. To demonstrate its utility, we applied this pipeline to analyze 32 blood biomarkers in 6,245 two-way AFR-EUR admixed individuals from the UK Biobank. This pipeline ran efficiently at scale, replicated known associations, and identified novel ancestry-specific loci. These novel associations were largely driven by variants present on African ancestral tracts but absent from European tracts, underscoring the value of local ancestry-aware methods in uncovering previously missed genetic signals. By enabling the efficient analysis of admixed individuals, our workflow facilitates Tractor use, paving the way for more broader genetic discovery.

genomics↗

Improved Allele Frequencies in gnomAD through Local Ancestry Inference

The Genome Aggregation Database (gnomAD) is a foundational resource for allele frequency data, widely used in genomic research and clinical interpretation. However, traditional estimates rely on individual-level genetic ancestry groupings that may obscure variation in recently admixed populations. To improve resolution, we applied local ancestry inference (LAI) to over 27 million variants in two admixed groups: Admixed American (n = 7,612) and African/African American (n = 20,250), deriving ancestry-specific allele frequencies. We show that 78.5% and 85.1% of variants in these groups, respectively, exhibit at least a twofold difference in ancestry-specific frequencies. Moreover, 81.49% of variants with LAI information would be assigned a higher gnomAD-wide maximum frequency after incorporating LAI, potentially altering clinical interpretations. This LAI-informed release reveals clinically relevant frequency differences that are masked in aggregate estimates and may support reclassifying some variants from Uncertain Significance to Benign or Likely Benign.

genomics↗

Characterizing features affecting local ancestry inference performance in admixed populations

In recent years, significant efforts have been made to improve methods for genomic studies of admixed populations using Local Ancestry Inference (LAI). Accurate LAI is crucial to ensure downstream analyses reflect the genetic ancestry of research participants accurately. Here, we test analytic strategies for LAI to provide guidelines for optimal accuracy, focusing on admixed populations reflective of Latin Americas primary continental ancestries - African (AFR), Amerindigenous (AMR), and European (EUR). Simulating LD-informed admixed haplotypes under a variety of 2 and 3-way admixture models, we implemented a standard LAI pipeline, testing three reference panel compositions to quantify their overall and ancestry-specific accuracy. We examined LAI miscall frequencies and true positive rates (TPR) across simulation models and continental ancestries. AMR tracts have notably reduced LAI accuracy as compared to EUR and AFR tracts in all comparisons, with TPR means for AMR ranging from 88-94%, EUR from 96-99% and AFR 98-99%. When LAI miscalls occurred, they most frequently erroneously called European ancestry in true Amerindigenous sites. Using a reference panel well-matched to the target population, even with a lower sample size, LAI produced true-positive estimates that were not statistically different from a high sample size but mismatched reference, while being more computationally efficient. While directly responsive to admixed Latin American cohort compositions, these trends are broadly useful for informing best practices for LAI across other admixed populations. Our findings reinforce the need for inclusion of more underrepresented populations in sequencing efforts to improve reference panels.

genomics↗