bioRxiv · 10.1101/2020.04.21.053876
XGMix: Local-Ancestry Inference With Stacked XGBoost
Abstract
AO_SCPLOWBSTRACTC_SCPLOWGenomic medicine promises increased resolution for accurate diagnosis, for personalized treatment, and for identification of population-wide health burdens at rapidly decreasing cost (with a genotype now cheaper than an MRI and dropping). The benefits of this emerging form of affordable, data-driven medicine will accrue predominantly to those populations whose genetic associations have been mapped, so it is of increasing concern that over 80% of such genome-wide association studies (GWAS) have been conducted solely within individuals of European ancestry [1]. The severe under-representation of the majority of the worlds populations in genetic association studies stems in part from an addressable algorithmic weakness: lack of simple, accurate, and easily trained methods for identifying and annotating ancestry along the genome (local ancestry). Here we present such a method (XGMix) based on gradient boosted trees, which, while being accurate, is also simple to use, and fast to train, taking minutes on consumer-level laptops.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Kumar, A., Montserrat, D. M., Bustamante, C., Ioannidis, A.. 2020-04-24. XGMix: Local-Ancestry Inference With Stacked XGBoost. https://doi.org/10.1101/2020.04.21.053876
Cite the original work for its findings. Save a collection to share your selection of sources.