Search bioRxiv⌕ Search

Biology subjects

Lipovac, J.

Publications and source records attributed to Lipovac, J..

4 recordsLinked to original sources

A complete diploid human genome benchmark for personalized genomics

Human genome resequencing typically involves mapping reads to a reference genome to call variants; however, this approach suffers from both technical and reference biases, leaving many duplicated and structurally polymorphic regions of the genome unmapped. Consequently, existing variant benchmarks, generated by the same methods, fail to assess these complex regions. To address this limitation, we present a telomere-to-telomere genome benchmark that achieves near-perfect accuracy (i.e. no detectable errors) across 99.4% of the complete, diploid HG002 genome. This benchmark adds 701.4 Mb of autosomal sequence and both sex chromosomes (216.8 Mb), totaling 15.3% of the genome that was absent from prior benchmarks. We also provide a diploid annotation of genes, transposable elements, segmental duplications, and satellite repeats, including 39,144 protein-coding genes across both haplotypes. To facilitate application of the benchmark, we developed tools for measuring the accuracy of sequencing reads, phased variant call sets, and genome assemblies against a diploid reference. Genome-wide analyses show that state-of-the-art de novo assembly methods resolve 2-7% more sequence and outperform variant calling accuracy by an order of magnitude, yielding just one error per 100 kb across 99.9% of the benchmark regions. Adoption of genome-based benchmarking is expected to accelerate the development of cost-effective methods for complete genome sequencing, expanding the reach of genomic medicine to the entire genome and enabling a new era of personalized genomics.

genomics↗

A Complete Telomere-to-Telomere Diploid Reference Genome for Indian Population

Human reference genomes have been instrumental in advancing genomic and biomedical research, but South and Southeast Asian populations are underrepresented, despite accounting for a large proportion of world population. As a part of effort on generating reference genomes for these populations, we present the first gapless, telomere-to-telomere (T2T) diploid genome assembly created by using a trio sample set of Indian ancestry (I002C), with NG50 of 154.89 Mb and 146.27 Mb for the maternal and paternal haplotypes, including the fully assembled rDNA array for the maternal chromosome 21 and Y chromosome. With the Merqury QVs of 82.05, 83.08 and 82.64 for the maternal, paternal and haploid assemblies respectively, I002C represents the highest-quality human genome assembled in both diploid and haploid forms to date. Compared to CHM13, the I002C genome displays substantial sequence diversity, resulting in 14,943 structural variants, including 3,236 novel variants absent from public databases. Analysis of trio-phased haplotypes further revealed elevated inter-haplotype divergence within centromeric and subtelomeric regions, along with identification of differentially methylated regions (DMRs) as candidates for novel imprinting loci. As a result of substantial SVs between them, I002C is a more suitable reference than CHM13 for the genomic analysis of South Asian samples with less reference bias and better performance in mapping and variant calling, particularly for long read sequencing data. As the first high-quality T2T diploid reference genome for Indian, the largest worlds population, I002C contributes to the growing set of population-specific reference genomes and helps to overcome a significant gap in human genome diversity.

genomics↗

MADRe: Strain-Level Metagenomic Classification Through Assembly-Driven Database Reduction

AbstractStrain-level metagenomic classification is essential for understanding microbial diversity and functional potential, but remains challenging, par- ticularly in the absence of prior knowledge about the composition of the sample. In this paper we present MADRe, a modular and scalable pipeline for long-read strain-level metagenomic classification, enhanced with Metagenome Assembly-Driven Database Reduction. MADRe com- bines long-read metagenome assembly, contig-to-reference mapping reas- signment based on an expectation-maximization algorithm for database reduction, and probabilistic read mapping reassignment to achieve sensi- tive and precise classification. We extensively evaluated MADRe on sim- ulated datasets, mock communities, and a real anaerobic digester sludge metagenome, demonstrating that it consistently outperforms existing tools by achieving higher precision with reduced false positives. MADRes de- sign allows users to apply either the database reduction or read classi- fication step individually. Using only the read classification step shows results on par with other tested tools. MADRe is open source and pub- licly available at https://github.com/lbcb-sci/MADRe.

bioinformatics↗

The Hitchhiker's Guide to Sequencing Data Types and Volumes for Population-Scale Pangenome Construction

Long-read (LR) technologies from Pacific Biosciences (PacBio) and Oxford Nanopore Technologies (ONT) have transformed genomics research by providing diverse data types like HiFi, Duplex, and ultra-long ONT (ULONT). Despite recent strides in achieving haplotype-phased gapless genome assemblies using long-read technologies, concerns persist regarding the representation of genetic diversity, prompting the development of pangenome references. However, pangenome studies face challenges related to data types, volumes, and cost considerations for each assembled genome, while striving to maintain sensitivity. The absence of comprehensive guidance on optimal data selection exacerbates these challenges. To fill this gap, our study evaluates available data types, their significance, and the required volumes for robust de novo assembly in population-level pangenome projects. The results show that achieving chromosome-level haplotype-resolved assembly requires 20x high-quality long reads (HQLR) such as PacBio HiFi or ONT duplex, combined with 15-20x of ULONT per haplotype and 30x of long-range data such as Omni-C. High-quality long reads from both platforms yield assemblies with comparable contiguity, with HiFi excelling in NG50 and phasing accuracies, while usage of duplex generates more T2T contigs. As Long-Read Technologies advance, our study reevaluates recommended data types and volumes, providing practical guidelines for selecting sequencing platforms and coverage. These insights aim to be vital to the pangenome research community, contributing to their efforts and pushing genomic studies with broader impacts.

bioinformatics↗