Search bioRxivSearch

Biology subjects

Halagan, M.

Publications and source records attributed to Halagan, M..

2 recordsLinked to original sources

Single haplotype admixture models using large scale HLA genotype frequencies to reproduce human admixture

The Human Leukocyte Antigen (HLA) is the most polymorphic region in humans. Anthropologists use HLA to trace populations migration and evolution. However, recent admixture between populations masks the ancestral haplotype frequency distribution.\n\nWe present an HLA-based method based on high-resolution HLA haplotype frequencies to resolve population admixture using a non-negative matrix factorization formalism and validated using haplotype frequencies from 56 populations. The result is a minimal set of original populations decoding roughly 90% of the total variance in the studied admixtures. These original populations agree with the geographical distribution, phylogenies and recent admixture events of the studied groups.\n\nWith the growing population of multi-ethnic individuals, the matching process for stem-cell and solid organ transplants is becoming more challenging. The presented algorithm provides a framework that facilitates the breakdown of highly admixed populations into original groups, which can be used to better match the rapidly growing population of multi-ethnic individuals worldwide.\n\nAuthor SummaryHuman Leukocyte Antigen (HLA) is known to be the most polymorphic region in the human genome. Anthropologists frequently use HLA to trace migration and evolution of different populations. This is due to the high linkage among HLA genes leading to the transmission of intact haplotypes from parents to offspring, hence preserving key population ancestral features.\n\nWe developed a new HLA-based method to identify admixture models in mixed populations using high-resolution HLA haplotype frequencies. Our results highlight that a single highly polymorphic locus can contain enough information to map clearly human admixture and the population genetics of the different human populations, and reproduces results based on SNP arrays.\n\nThe presented algorithm is validated using haplotype frequencies sampled from 56 worldwide populations. Under such factorization we demonstrate that 90% of the variance in these populations can be explained using a much-reduced set of 8 ethnic groups. We demonstrate that the estimated ethnic groups and admixture models agree with the geographical distribution, population phylogenies and recent historic admixture events of the studied populations.

immunology

GRIMM: GRaph IMputation and Matching for HLA Genotypes

Motivation: For over 10 years allele-level HLA matching for bone marrow registries has been performed in a probabilistic context. HLA typing technologies provide ambiguous results in that they could not distinguish among all known HLA allele sequences, therefore registries have implemented matching algorithms that provide lists of donor and cord blood units ordered in terms of the likelihood of allele-level matching at specific HLA loci. With the growth of registry sizes, current match algorithm implementations are unable to provide match results in real time.\n\nResults: We present here novel computationally-efficient open source implementation of an HLA imputation and match algorithm using a graph database platform. Using graph traversal, our algorithm runtime grows slowly with registry size. This implementation generates results that agree with consensus output on a publicly-available match algorithm crossvalidation dataset.\n\nAvailability: The Python, Perl and Neo4jJcode is available at https://git.com/nmdp-bioinformatics/grimm\n\nSupplementary information: Supplementary data are available at Bioinformatics online.

bioinformatics