Search bioRxiv⌕ Search

Biology subjects

Doroschuk, N.

Publications and source records attributed to Doroschuk, N..

3 recordsLinked to original sources

HLA alleles and haplotype distribution across Russian population groups

HLA loci are highly polymorphic genome regions, with allele frequencies varying significantly across different populations. Population HLA frequency databases may contain biases and make cross-study comparison complicated due to varying data curation protocols, genotyping methodologies, resolution, and inconsistencies in the selection criteria for population samples. This study presents HLA allele frequencies of class I (HLA-A, -B, -C) and class II (HLA-DRB1, -DQB1, -DQA1) as well as their combined haplotypes obtained from over 18,000 whole genome sequencing samples of the Russian population. Cohort was stratified based on PCA and admixture components providing frequencies for 14 different ethnic groups. For 12 groups cohort size allowed us to reach average saturation of 96% of allele frequencies in groups. Moreover, we demonstrated the utility of composed statistics for disease populational study using type 1 diabetes (T1D) as an example. Populations with similar aggregated genetic risk for T1D demonstrated substantial differences in frequencies of risk and protective HLA alleles. Obtained frequency data was made publicly available through the Allele Frequency Net Database improving previously sparse coverage in HLA frequencies data for east Europe and north Asia regions.

genetics↗

Systematic analysis of insertions signature in gnomAD revealed large set of novel processed pseudogenes

Pseudogenes are non-functional copies of protein-coding genes that arise through genomic duplication or retrotransposition. Processed pseudogenes (PPs) is the most abundant class of pseudogenes, which is generated via mRNA reverse transcription and subsequent cDNA integration. Presence of PPs complicates the analysis of short read sequencing data due to high similarity with parental gene and frequent absence from reference genome. Here we demonstrate that the presence of non-reference (absent from reference genome) PPs leads to the very distinctive artefact of germline variant calling - long insertions on exon-intron boundaries, which sequences could be mapped to other exons of the same gene. We showed that by detecting these artifacts it is possible to identify non-reference PPs existence based on the cohort summary statistics without analysing sample-level data. We used identified signature of PPs presence to systematically mine the gnomAD database which currently contains over 70,000 whole-genome and over 700,000 exome samples to describe novel non-reference PPs. Our approach uncovered 1498 non-reference PPs of which 1268 were novel and absent in the latest GENCODE release. This resource enhances the accuracy of variant interpretation and contributes to a deeper understanding of pseudogenes diversity across human populations.

genomics↗

Sanger validation of WGS variants - when to?

With the development of Next-Generation Sequencing (NGS) technologies it became possible to simultaneously analyze millions of variants. Despite the quality improvement it is generally still required to confirm the variants before reporting. However, in recent years the dominant idea is that one could define the quality thresholds for "high quality" variants which do not require orthogonal validation. Despite that, no works to date report the concordance between variants from whole genome sequencing and their gold-standard Sanger validation. In this study we analyzed the concordance for 1756 WGS variants in order to establish the appropriate thresholds for high-quality variants filtering. Resulting thresholds allowed us to drastically reduce the number of variants which require validation, to 5,6% and 1.2% of the initial set for caller-agnostic thresholds and caller-dependent QUAL threshold respectively.

genomics↗