Search bioRxiv⌕ Search

Biology subjects

Zhou, X. M.

Publications and source records attributed to Zhou, X. M..

10 recordsLinked to original sources

Enhancing variant detection in complex genomes: leveraging linked reads for robust SNP, Indel, and structural variant analysis

Accurate detection of genetic variants, including single nucleotide polymorphisms (SNPs), small insertions and deletions (INDELs), and structural variants (SVs), is essential for comprehensive genomic analysis. While short-read sequencing performs well for SNP and INDEL detection, it remains limited in resolving SVs, particularly in complex genomic regions, due to its short read length. Linked-read sequencing technologies, such as single-tube Long Fragment Read (stLFR), partially address this limitation by incorporating molecular barcodes to provide long-range information. In this study, we evaluate conventional paired-end linked reads (PE100_stLFR) and explore a conceptual extension: long single-end barcoded reads of 500 bp (SE500_stLFR) and 1000 bp (SE1000_stLFR). We developed_stLFR-sim, a Python-based simulator that reproduces the_stLFR workflow and enables realistic benchmarking. Using a high-quality T2T assembly of HG002, we generated multiple datasets across 12 sequencing configurations. SVs were called using Aquila_stLFR (v2) and benchmarked against the Genome in a Bottle (GIAB) HG002 SV truth set with Truvari. We show that simulated PE100_stLFR closely matches real data, validating the simulation framework. Increasing read length consistently improves SV detection accuracy, with SE1000_stLFR achieving the best performance and approaching long-read methods while outperforming short-read and pangenome-based approaches. Collectively, our results highlight the strong potential of long single-end barcoded reads for improving SV detection, and suggest that even modest increases in read length, when combined with barcode information, can provide a cost-effective and practical strategy for enhancing future sequencing technologies and SV discovery.

bioinformatics↗

stDyer-image improves clustering analysis of spatially resolved transcriptomics and proteomics with morphological images

Spatially resolved transcriptomics (SRT) and spatially resolved proteomics (SRP) data enable the study of gene expression and protein abundances within their precise spatial and cellular contexts in tissues. Certain SRT and SRP tech-nologies also capture corresponding morphology images, adding another layer of valuable information. However, few existing methods developed for SRT data effectively leverage these supplementary images to enhance clustering performance. Here, we introduce stDyer-image, an end-to-end deep learning framework designed for clustering for SRT and SRP datasets with images. Unlike existing methods that utilize images to complement gene expression data, stDyer-image directly links image features to cluster labels. This approach draws inspiration from pathologists, who can visually identify specific cell types or tumor regions from morphological images without relying on gene expression or protein abundances. Benchmarks against state-of-the-art tools demonstrate that stDyer-image achieves superior performance in clustering. Moreover, it is capable of handling large-scale datasets across diverse technologies, making it a versatile and powerful tool for spatial omics analysis.

bioinformatics↗

SCGclust: Single Cell Graph clustering using graphautoencoders integrating SNVs and CNAs

Intra-tumor heterogeneity (ITH) is a compounding factor for cancer prognosis and treatment. Single-cell DNA sequencing (scDNA-seq) provides cellular resolution of the variations in a cell and has been widely used to study cancer progression and responses to drug and treatment. While the low coverage scDNA-seq technologies typically provides a large number of cells, accurate cell clustering is essential for effectively characterizing ITH. Existing cell clustering methods typically are based on either single nucleotide variations (SNV) or copy number alterations (CNA), without leveraging both signals together. Since both SNVs and CNAs are indicative of the cell subclonality, in this paper, we designed a robust cell clustering tool that integrates both signals using a graph autoencoder. Our model co-trains the graph autoencoder and a graph convolutional network (GCN) to guanrantee meaningful clustering results and to prevent all cells from collapsing into a single cluster. Given the low dimensional embedding generated by the autoencoder, we adopted a Gaussian Mixture Model to further cluster cells. We evaluated our method on eight simulated datasets and a real cancer sample. Our results demonstrate that our method consistently achieves higher V-measure scores compared to SBMClone, a SNV-based method, and a K-means method, which relies solely on CNA signals. These findings highlight the advantage of integrating both SNV and CNA signals within a graph autoencoder framework for accurate cell clustering. SCGclust is publicly available at https://github.com/compbio-mallory/cellClustering_GNN.

bioinformatics↗

FocalSV: target region-based structural variant assembly and refinement using single-molecule long read sequencing data

Structural variants (SVs) play a critical role in shaping the diversity of the human genome and their detection holds significant potential for advancing precision medicine. Despite notable progress in single-molecule long-read sequencing technologies, accurately identifying SV breakpoints and resolving their sequence remains a major challenge. Current alignment-based tools often struggle with precise breakpoint detection and sequence characterization, while whole genome assembly-based methods are computationally demanding and less practical for targeted analyses. Neither approach is ideally suited for scenarios where regions of interest are predefined and require precise SV characterization. To address this gap, we introduce FocalSV, a target region assembly-based SV detection tool that combines the precision of assembly-based methods with the efficiency of region-specific approaches. FocalSV was evaluated on nine germline datasets and two paired normal-tumor cancer datasets, demonstrating superior performance in both precision and efficiency.

bioinformatics↗

Brain structure and activity predicting cognitive maturation in adolescence

Cognitive abilities of primates, including humans, continue to improve through adolescence 1,2. While a range of changes in brain structure and connectivity have been documented 3,4, how they affect neuronal activity that ultimately determines performance of cognitive functions remains unknown. Here, we conducted a multilevel longitudinal study of monkey adolescent neurocognitive development. The developmental trajectory of neural activity in the prefrontal cortex accounted remarkably well for working memory improvements. While complex aspects of activity changed progressively during adolescence, such as the rotation of stimulus representation in multidimensional neuronal space, which has been implicated in cognitive flexibility, even simpler attributes, such as the baseline firing rate in the period preceding a stimulus appearance had predictive power over behavior. Unexpectedly, decreases in brain volume and thickness, which are widely thought to underlie cognitive changes in humans 5 did not predict well the trajectory of neural activity or cognitive performance changes. Whole brain cortical volume in particular, exhibited an increase and reached a local maximum in late adolescence, at a time of rapid behavioral improvement. Maturation of long-distance white matter tracts linking the frontal lobe with areas of the association cortex and subcortical regions best predicted changes in neuronal activity and behavior. Our results provide evidence that optimization of neural activity depending on widely distributed circuitry effects cognitive development in adolescence.

neuroscience↗

Leveraging cross-source heterogeneity to improve the performance of bulk gene expression deconvolution

We introduce CSsingle, a novel method that enhances the decomposition of bulk and spatial transcriptomic (ST) data by addressing key challenges in cellular heterogeneity. CSsingle applies cell size correction using ERCC spike-in controls, enabling it to account for variations in RNA content between cell types and achieve accurate bulk data deconvolution. In addition, it enables fine-scale analysis for ST data, advancing our understanding of tissue architecture and cellular interactions, particularly in complex microenvironments. We provide a unified tool for integrating bulk and ST with scRNA-seq data, advancing the study of complex biological systems and disease processes. The benchmark results demonstrate that CSsingle outperforms existing methods in accuracy and robustness. Validation using more than 700 normal and diseased samples from gastroesophageal tissue reveals the predominant presence of mosaic columnar cells (MCCs), which exhibit a gastric and intestinal mosaic phenotype in Barretts esophagus and esophageal adenocarcinoma (EAC), in contrast to their very low detectable levels in esophageal squamous cell carcinoma and normal gastroesophageal tissue. We revealed a dynamic relationship between MCCs and squamous cells during immune checkpoint inhibitors (ICI)-based treatment in EAC patients, suggesting MCC expression signatures as predictive and prognostic markers of immunochemotherapy outcomes. Our findings reveal the critical role of MCC in the treatment of EAC and its potential as a biomarker to predict outcomes of immunochemotherapy, providing insight into tumor epithelial plasticity to guide personalized immunotherapeutic strategies.

bioinformatics↗

Benchmarking clustering, alignment, and integration methods for spatial transcriptomics

Spatial transcriptomics (ST) is advancing our understanding of complex tissues and organisms. However, building a robust clustering algorithm to define spatially coherent regions in a single tissue slice, and aligning or integrating multiple tissue slices originating from diverse sources for essential downstream analyses remain challenging. Numerous clustering, alignment, and integration methods have been specifically designed for ST data by leveraging its spatial information. The absence of benchmark studies complicates the selection of methods and future method development. Here we systematically benchmark a variety of state-of-the-art algorithms with a wide range of real and simulated datasets of varying sizes, technologies, species, and complexity. Different experimental metrics and analyses, like adjusted rand index (ARI), uniform manifold approximation and projection (UMAP) visualization, layer-wise and spot-to-spot alignment accuracy, spatial coherence score (SCS), and 3D reconstruction, are meticulously designed to assess method performance as well as data quality. We analyze the strengths and weaknesses of each method using diverse quantitative and qualitative metrics. This analysis leads to a comprehensive recommendation that covers multiple aspects for users. The code used for evaluation is available on GitHub. Additionally, we provide jupyter notebook tutorials and documentation to facilitate the reproduction of all benchmarking results and to support the study of new methods and new datasets (https://benchmarkst-reproducibility.readthedocs.io/en/latest/).

bioinformatics↗

MaskGraphene: Advancing joint embedding, clustering, and batch correction for spatial transcriptomics using graph-based self-supervised learning

Recent advancements in spatial transcriptomics (ST) have underscored the importance of integrating data from multiple ST slices for joint analysis. A major challenge remains generating interpretable joint embeddings that preserve geometric information for downstream analyses. Here we introduce MaskGraphene, a graph neural network that combines self-supervised and self-contrastive training to integrate gene expression and spatial location into joint embeddings. By employing clusterwise alignment and a graph attention autoencoder with masked self-supervised and triplet loss optimizations, MaskGraphene effectively preserves geometric structures while achieving batch correction. In benchmarks against seven state-of-the-art methods, MaskGraphene consistently demonstrated superior alignment accuracy and geometric fidelity across diverse ST datasets. Its interpretable embeddings significantly enhanced downstream applications, including domain identification, spatial trajectory reconstruction, biomarker discovery, and the creation of topographical maps of brain slices. Notably, MaskGraphene successfully recovered layer-wise brain structures with near-perfect accuracy. MaskGraphene provides a powerful and versatile framework for advancing ST data integration and analysis, unlocking valuable biological insights.

bioinformatics↗

CNVeil enables accurate and robust tumor subclone identification and copy number estimation from single-cell DNA sequencing data

Single-cell DNA sequencing (scDNA-seq) has significantly advanced cancer research by enabling precise detection of chromosomal aberrations, such as copy number variations (CNVs), at a single-cell level. These variations are crucial for understanding tumor progression and heterogeneity among tumor subclones. However, accurate CNV inference in scDNA-seq has been constrained by several factors, including low coverage, sequencing errors, and data variability. To address these challenges, we introduce CNVeil, a robust quantitative algorithm designed to accurately reveal CNV profiles while overcoming the inherent noise and bias in scDNA-seq data. CNVeil incorporates a unique bias correction method using normal cell profiles identified by a PCA-based Gini coefficient, effectively mitigating sequencing bias. Subsequently, a multi-level hierarchical clustering, based on selected highly variable bins, is employed to initially identify coarse subclones for robust ploidy estimation and further identify fine subclones for segmentation. To infer the CNV segmentation landscape, a novel change rate-based across-cell breakpoint identification approach is specifically designed to diminish the effects of low coverage and data variability on a per-cell basis. Finally, a consensus segmentation is utilized to further standardize read depth for the inference of the final CNV profile. In comprehensive benchmarking experiments, where we compared CNVeil with seven state-of-the-art CNV detection tools, CNVeil exhibited exceptional performance across a diverse set of simulated and real scDNA-seq data in cancer genomics. CNVeil excelled in subclone identification, segmentation, and CNV profiling. In light of these results, we anticipate that CNVeil will significantly contribute to single-cell CNV analysis, offering enhanced insights into chromosomal aberrations and genomic complexity.

bioinformatics↗

Laminar pattern of adolescent development changes in working memory neuronal activity

Adolescent development is characterized by an improvement in cognitive abilities, such as working memory. Neurophysiological recordings in a non-human primate model of adolescence have revealed changes in neural activity that mirror improvement in behavior, including higher firing rate during the delay intervals of working memory tasks. The laminar distribution of these changes is unknown. By some accounts, persistent activity is more pronounced in superficial layers, so we sought to determine whether changes are most pronounced there. We therefore analyzed neurophysiological recordings from neurons recorded in the young and adult stage, at different cortical depths. Superficial layers exhibited increased baseline firing rate in the adult stage. Unexpectedly, changes in persistent activity were most pronounced in the middle layers. Finally, improved discriminability of stimulus location was most evident in the deeper layers. These results reveal the laminar pattern of neural activity maturation that is associated with cognitive improvement. NEW AND NOTEWORTHYStructural brain changes are evident during adolescent development particularly in the cortical thickness of the prefrontal cortex, at a time when working memory ability increases markedly. The depth distribution of neurophysiological changes during adolescence is not known. Here we show that neurophysiological changes are not confined to superficial layers, which have most often been implicated in the maintenance of working memory. Contrary to expectations, greatest changes were evident in intermediate layers of the prefrontal cortex.

neuroscience↗