Search bioRxiv⌕ Search

Biology subjects

Yakymenko, I.

Publications and source records attributed to Yakymenko, I..

2 recordsLinked to original sources

Accurate imputation of inversions in human genomes using different algorithms and data sources

Complex genomic regions harbor different structural arrangements that can mutate quite rapidly, which makes determining their functional effects very difficult. Characterization of inversions originated by homologous mechanisms is especially challenging due to the presence of inverted repeats at the breakpoints and the fact that most of them are recurrent. Imputation can infer missing genotypes, but it has been mainly limited to simple variants and little is known about how well it works for human inversions. Here, we tested five common imputation programs to impute a set of 52 inversions experimentally genotyped in multiple samples that lacked SNPs in perfect linkage disequilibrium. Using whole genome sequencing data and simulated microarrays with variable SNP density, we found that 40.4-75.5% of inversions could be accurately imputed in three human populations by at least one program, with results depending mainly on the number of SNPs available, the genotyped samples and the recurrence of inversions. Also, genotype probability filtering was a key factor for inversion imputation accuracy. In particular, Minimac4 and IMPUTE5 showed more accurately imputed inversions and less poorly imputed individuals with respect to the other methods. This work therefore contributes to optimize inversion imputation, making possible the study of their functional impact.

genomics↗

Resolving the full set of human polymorphic inversions and other complex variants from ultra-long read data

Inversions are a unique type of balanced structural variants (SVs) with important consequences in multiple organisms. However, despite considerable effort, this and other complex SVs remain poorly characterized due to the presence of large repeats. New techniques are finally allowing us to identify the full spectrum of human inversions, but the number of individuals analyzed is still quite limited. Here, we take advantage of Oxford Nanopore Technologies (ONT) long reads to characterize an exhaustive catalogue of 612 candidate inversions between 197 bp and 4.4 Mb of length and flanked by <190-kb long inverted repeats (IRs). For that, we developed a bioinformatic package to identify inversion alleles reliably from long read data. Next, using a combination of different DNA extraction, library preparation, and ONT sequencing protocols, we showed that ultra-long reads (50-100 kb) and adaptive sampling are an efficient method to detect most human inversions. Lastly, by analyzing ONT data from 54 diverse individuals, 87-99% of the inversions could be genotyped in each sample, depending mainly on read and IR length and genome coverage. Both orientations were observed for 155 of the analyzed regions (frequency 0.01-0.49), which multiplies by three the polymorphic IR-mediated inversions studied in detail so far. Moreover, we found more than 300 additional independent SVs in the studied regions and resolved several complex rearrangements. Our work therefore provides an accurate benchmark of those inversions that typically escape most analyses, improving existing resources, such as the Pangenome. In addition, it demonstrates the potential of nanopore sequencing to determine the functional impact of missing human genomic variation.

genomics↗