Search bioRxiv⌕ Search

Biology subjects

Gambin, T. L.

Publications and source records attributed to Gambin, T. L..

2 recordsLinked to original sources

The Thousand Polish Genomes Project- a national database of Polish variant allele frequencies

Although Slavic populations account for over 3.5% of world inhabitants, no centralized, open source reference database of genetic variation of any Slavic population exists to date. Such data are crucial for either biomedical research and genetic counseling and are essential for archeological and historical studies. Polish population, homogenous and sedentary in its nature but influenced by many migrations of the past, is unique and could serve as a good genetic reference for middle European Slavic nations. The aim of the present study was to describe first results of analyses of a newly created national database of Polish genomic variant allele frequencies. Never before has any study on the whole genomes of Polish population been conducted on such a large number of individuals (1,079). A wide spectrum of genomic variation was identified and genotyped, such as small and structural variants, runs of homozygosity, mitochondrial haplogroups and Mendelian inconsistencies. The allele frequencies were calculated for 943 unrelated individuals and released publicly as The Thousand Polish Genomes database. A precise detection and characterisation of rare variants enriched in the Polish population allowed to confirm the allele frequencies for known pathogenic variants in diseases, such as Smith-Lemli-Opitz syndrome (SLOS) or Nijmegen breakage syndrome (NBS). Additionally, the analysis of OMIM AR genes led to the identification of 22 genes with significantly different cumulative allele frequencies in the Polish (POL) vs European NFE population. We hope that The Thousand Polish Genomes database will contribute to the worldwide genomic data resources for researchers and clinicians.

genomics↗

De novo mutation in ancestral generations evolves haplotypes contributing to disease

PurposeThe variome of the Turkish (TK) population, a population with a considerable history of admixture and consanguinity, has not been deeply investigated deeply for its potential impact on the genomic architecture of disease traits. MethodsWe generated and analyzed a database of variants derived from exome sequencing (ES) data of 773 TK unrelated, clinically affected individuals with various suspected Mendelian disease traits, and 643 unaffected relatives. ResultsUsing uniform manifold approximation and projection (UMAP), we showed that the TK genomes are more similar to those of Europeans and consist of two main subpopulations: clusters 1 and 2 (N=235 and 1,181) that differ in admixture proportion and variome (https://turkishvariomedb.shinyapps.io/tvdb/). Furthermore, the higher inbreeding coefficient (F) values observed in the TK affected compared to unaffected individuals correlated with a larger median span of long-sized (>2.64 Mb) runs of homozygosity (ROH) regions (p-value=2.09e-18). We show that long-sized ROHs are more likely to be formed on recently configured haplotypes enriched for rare homozygous deleterious variants in the TK-affected compared to TK-unaffected individuals (p-value= 3.35e-11). Analysis of genotype-phenotype correlations reveals that genes with rare homozygous deleterious variants in long-sized ROHs provide the most comprehensive set of molecular diagnoses for the observed disease traits with a systematic quantitative analysis of HPO (Human Phenotype Ontology) terms. ConclusionOur findings support the notion that novel rare variants on newly configured haplotypes arising within the recent past generations of a family or clan contribute significantly to recessive disease traits in the TK population.

genomics↗