Search bioRxiv⌕ Search

Biology subjects

Bhamidipati, S. V.

Publications and source records attributed to Bhamidipati, S. V..

5 recordsLinked to original sources

Somatic variant detection in normal tissues from single-cell sequencing data

A crucial advantage of single-cell sequencing (SCS) is its ability to identify somatic variants in individual cells, enabling phylogenetic analysis of cellular populations within bulk tissues. While identifying somatic variants in tumor tissues via SCS has become a common practice, doing so in normal tissues remains challenging due to the rarity of somatic variants in normal cells. To evaluate the feasibility of somatic variant calling from widely available single-nucleus RNA-seq (snRNA-seq) and single-nucleus ATAC-seq (snATAC-seq) data, we profiled a Cell-line mix of six HapMap samples prepared by the SMaHT consortium using 10x Genomics 5 snRNA-seq (12k cells with 36k mean reads per cell) and snATAC-seq (11k cells with 14k median high-quality fragments per cell) for variant calling. PacBio long-read whole genome sequencing (WGS) data (109x) generated from individual cell lines were used as ground truth. Two computational tools, Monopogen and SComatic, were used for somatic variant calling from the SCS data. Monopogen achieved single nucleotide variant (SNV) detection accuracies of 93.30% in the snRNA-seq and 99.64% in the snATAC-seq data, both of which outperformed SComatic (74.35% and 94.29%, respectively). Monopogen also consistently detected somatic SNVs at cellular fractions as low as 0.5% (2.54% in snRNA and 0.81% in snATAC) in individual samples. Notably, snATAC-seq exhibited higher genomic coverage breadth and larger number of variants detected than snRNA-seq. While the SCS data have lower overall genome coverage than that of the bulk WGS, the single-cell level variant resolution allows Monopogen to assign variants to their cells of origin with over 80% accuracy in both RNA and ATAC modalities, thereby facilitating studies of clonal evolution and cell-type-specific mutagenesis. Other benchmarking methods were also evaluated (DeepVariant, Cellsnp-lite and Mutect2) for comparison. In conclusion, our study demonstrated the feasibility of performing reliable single-cell somatic mutation calling in a cell-line mixture and discussed the strengths and limitations of current computational methods when applied to normal tissues.

bioinformatics↗

Expanding the Genome in a Bottle Truth Set: Detection and Validation of Novel Low-frequency Variants Using High-accuracy NanoSeq

HighlightsO_LINanoSeq-MBN achieves near-genome, Poisson-like coverage with minimal trinucleotide bias. C_LIO_LIExpands the GIAB truth set by up to 160k de novo variants. C_LIO_LIAdds a somatic layer to GIAB, enabling benchmarking and calibration of rare variants. C_LIO_LIHigh-CADD exonic and splice variants highlight value for surveillance and clinical triage. C_LI Somatic mutations record tissue molecular history and inform risk, prognosis, and therapy, yet their variant allele fractions often fall below the reliable detection limit of conventional short-read sequencing. In contrast, duplex sequencing technology featured by NanoSeq applies the principle of single molecule detection and thereby overcomes the limitation. However, the original NanoSeq protocol relies on the restriction enzyme-based genome fragmentation, which constrained its genome coverage to 30-40%. To enable whole-genome discovery with duplex-level fidelity, we pursued two complementary approaches to optimize the NanoSeq protocol: (i) a restriction-enzyme strategy densifies accessible sites using orthogonal 4-bp cutters; and (ii) a workflow using sonication followed by mung bean nuclease with T4 polynucleotide kinase, Klenow fragment and dATP/ddBTP mixture (NanoSeq-MBN) to blunt and repair/A-tailing DNA, while minimizing repair artifacts. We systematically benchmarked their performance using Genome in a Bottle (GIAB) gold-standard sample mixtures. As a result, NanoSeq-MBN achieved near genome-wide, Poissonlike coverage with minimal trinucleotide-context bias and ultra-high accuracy. Beyond variants already present in the GIAB truth set, NanoSeq-MBN identified approximately 120,000-160,000 de novo mutations per sample missing in the truth set, Notably, over 98% had orthogonal support in reanalyzed GIAB bulk Illumina HiSeq libraries. These novel variants extended GIAB from germline benchmarking to rare-variant discovery and calibration of subclonal detection. Functional annotation revealed enrichment of high Combined Annotation Dependent Depletion (CADD) scores mutations in exonic and splice-related regions. Variants intersecting ClinVar entries and OMIM genes highlighted potential for surveillance and clinical triage. Collectively, these results add a somatic layer to GIAB, enabling calibration of burdens and mutational signatures in lymphoblastoid lines and provide reference material for rare-variant assays. The NanoSeq-MBN workflow offers a path to whole-genome, high-fidelity discovery of ultra-rare somatic variation with relevance to clinical assay validation.

molecular biology↗

Pre-Sequencing Assessment of RNA-Seq Library Quality Using Real-Time qPCR

RNA sequencing (RNA-Seq) is an essential sequencing assay for studying transcriptome profiling. Ribosomal RNA (rRNA) comprises more than 80 - 90% of total cellular RNA, efficient rRNA removal is essential for accurately capturing the transcriptome, particularly to sequence low-abundance mRNAs. Inefficient rRNA removal during library preparation can result from variations in sample quality, preparation methods and handling. Estimating rRNA content in RNA-Seq libraries pre-sequencing is therefore challenging due to the absence of a reliable and cost-effective assessment method. This study addresses the issue by introducing a scalable qPCR-based assay targeting 18S rRNA to evaluate rRNA depletion efficiency pre-sequencing. qPCR efficiency was optimized using serial dilutions of Universal Human Reference (UHR) control and Ct thresholds were established using pilot data from 644 libraries. Following this optimization, analysis of 1,748 Total RNA-Seq libraries and 445 Poly A+ two widely used RNA-Seq library methods, demonstrated a strong correlation between 18S rRNA qPCR results and post-sequencing rRNA rates. This assay was also used to evaluate the performance of Oligo dT beads from four different vendors to enrich mRNA. This 18S rRNA qPCR assay is a cost-effective, scalable approach for reliably predicting rRNA read percentage in RNA-Seq libraries pre-sequencing. Method SummaryThe 18S rRNA-specific qPCR assay is a cost-effective method for evaluating rRNA depletion efficiency in RNA-Seq libraries prior to sequencing by targeting the 18S ribosomal RNA. qPCR efficiency was optimized, and Ct thresholds were set from pilot data and validated across 1,748 Total RNA and 445 Poly A+ RNA-Seq libraries. A Ct threshold of [≥]16 for Total RNA and Ct 13-14 or greater for Poly A+ libraries demonstrates a strong correlation between qPCR results and post-sequencing rRNA read percentages. The assay enables early identification of poorly depleted samples, reducing unnecessary sequencing costs. Additionally, the method was employed to benchmark oligo-dT beads from four vendors, showing its utility in evaluating RNA-seq library preparation kits. This scalable approach supports quality control in both research and high-throughput sequencing environments.

molecular biology↗

VizCNV: An integrated platform for concurrent phased BAF and CNV analysis with trio genome sequencing data

BackgroundCopy number variation (CNV) is a class of genomic Structural Variation (SV) that underlie genomic disorders and can have profound implications for health. Short-read genome sequencing (sr-GS) enables CNV calling for genomic intervals of variable size and across multiple phenotypes. However, unresolved challenges include an overwhelming number of false-positive calls due to systematic biases from non-uniform read coverage and collapsed calls resulting from the abundance of paralogous segments and repetitive elements in the human genome. MethodsTo address these interpretative challenges, we developed VizCNV. The VizCNV computational tool for inspecting CNV calls uses various data signal sources from sr-GS data, including read depth, phased B-allele frequency, as well as benchmarking signals from other SV calling methods. The interactive features and view modes are adept for analyzing both chromosomal abnormalities [e.g., aneuploidy, segmental aneusomy, and chromosome translocations], gene exonic CNV and non-coding gene regulatory regions. In addition, VizCNV includes a built-in filter schema for trio genomes, prioritizing the detection of impactful germline CNVs, such as de novo CNVs. Upon computational optimization by fine-tuning parameters to maximize sensitivity and specificity, VizCNV demonstrated approximately 83.8% recall and 77.2% precision on the 1000 Genome Project data with an average coverage read depth of 30x. ResultsWe applied VizCNV to 39 families with primary immunodeficiency disease without a molecular diagnosis. With implemented build-in filter, we identified two de novo CNVs and 90 inherited CNVs >10 kb per trio. Genotype-phenotype analyses revealed that a compound heterozygous combination of a paternal 12.8 kb deletion of exon 5 and a maternal missense variant allele of DOCK8 are likely the molecular cause of one proband. ConclusionsVizCNV provides a robust platform for genome-wide relevant CNV discovery and visualization of such CNV using sr-GS data.

genomics↗

Complete Genomic Characterization of Global Pathogens, Respiratory Syncytial Virus (RSV), and Human Norovirus (HuNoV) Using Probe-based Capture Enrichment.

Respiratory syncytial virus (RSV) is the leading cause of lower respiratory tract infections in children worldwide, while human noroviruses (HuNoV) are a leading cause of epidemic and sporadic acute gastroenteritis. Generating full-length genome sequences for these viruses is crucial for understanding viral diversity and tracking emerging variants. However, obtaining high-quality sequencing data is often challenging due to viral strain variability, quality, and low titers. Here, we present a set of comprehensive oligonucleotide probe sets designed from 1,570 RSV and 1,376 HuNoV isolate sequences in GenBank. Using these probe sets and a capture enrichment sequencing workflow, 85 RSV positive nasal swab samples and 55 (49 stool and six human intestinal enteroids) HuNoV positive samples encompassing major subtypes and genotypes were characterized. The Ct values of these samples ranged from 17.0-29.9 for RSV, and from 20.2-34.8 for HuNoV, with some HuNoV having below the detection limit. The mean percentage of post-processing reads mapped to viral genomes was 85.1% for RSV and 40.8% for HuNoV post-capture, compared to 0.08% and 1.15% in pre-capture libraries, respectively. Full-length genomes were>99% complete in all RSV positive samples and >96% complete in 47/55 HuNoV positive samples--a significant improvement over genome recovery from pre-capture libraries. RSV transcriptome (subgenomic mRNAs) sequences were also characterized from this data. Probe-based capture enrichment offers a comprehensive approach for RSV and HuNoV genome sequencing and monitoring emerging variants. IMPORTANCERespiratory syncytial virus (RSV) and human noroviruses (HuNoV) are NIAID category C and category B priority pathogens, respectively, that inflict significant health consequences on children, adults, immunocompromised patients, and the elderly. Due to the high strain diversity of RSV and HuNoV genomes, obtaining complete genomes to monitor viral evolution and pathogenesis is challenging. In this paper, we present the design, optimization, and benchmarking of a comprehensive oligonucleotide target capture method for these pathogens. All 85 RSV samples and 49/55 HuNoV samples were patient-derived with six human intestinal enteroids. The methodology described here results has a higher success rate in obtaining full-length RSV and HuNoV genomes, enhancing the efficiency of studying these viruses and mutations directly from patient-derived samples.

molecular biology↗