Search bioRxiv⌕ Search

Biology subjects

Singh, O.

Publications and source records attributed to Singh, O..

3 recordsLinked to original sources

Path2Omics: Enhanced transcriptomic and methylation prediction accuracy from tumor histopathology

Precision oncology is becoming increasingly integral to clinical practice, demonstrating notable improvements in treatment outcomes. While molecular data provide comprehensive insights, obtaining such data remains costly and time-consuming. To address this challenge, we developed Path2Omics, a deep learning model that predicts gene expression and methylation from histopathology for 23 cancer types. Path2Omics was trained on 20,497 slides (9,456 formalin-fixed and paraffin-embedded (FFPE) and 11,041 fresh frozen (FF)) from 8,007 patients across 23 The Cancer Genome Atlas cohorts. When tested on FFPE slides, the most readily available format in clinical pathology practice, the integrated model outperformed its individual FF and FFPE components, robustly predicting nearly 5,000 genes on average, approximately five times more than our recently published DeepPT model. Externally evaluated on seven independent cohorts, Path2Omics robustly predicted the expression of approximately 4,400 genes, yielding a 30% increase over the FFPE model alone. Finally, we demonstrate that the inferred gene expression is nearly as effective as the actual values in predicting patient survival and treatment response. These results lay the basis for using Path2Omics to advance precision oncology from histopathology slides in a speedy and cost-effective manner. Statement of significancePath2Omics is a deep learning model that accurately predicts gene expression and methylation from histopathology slides across 23 cancer types. Unlike existing approaches that rely solely on FFPE slides for training, Path2Omics leverages both FFPE and FF slides by constructing two separate models and integrating them. Downstream analyses show that the inferred values from Path2Omics are nearly as effective as actual values in predicting patient survival and treatment response.

bioinformatics↗

Bacterial ADP-heptose initiates a revival stem cell program in the intestinal epithelium

The intestinal epithelium has an exceptional capacity to repair following injury, and recent evidence has suggested that YAP-dependent signaling was crucial for the expansion of Clu+ revival stem cells (revSCs) with fetal-like characteristics, which are essential for epithelial regeneration. However, neither the mechanism underlying where these revSCs emerge from nor the nature of the physiological cues that induce this revSC program, are clearly identified. Here, we first demonstrate that Alpk1 and Tifa, which encode the proteins essential for the detection of the bacterial metabolite ADP-heptose (ADP-Hep), were expressed by the stem cell pool in the intestinal epithelium. Treatment of intestinal organoids with ADP-Hep not only induced acute NF-{kappa}B pro-inflammatory signaling but also TNF-dependent apoptosis within the crypt, causing blunted proliferation and acute disruption of the crypt architecture, while also triggering induction of a revSC program. To identify the molecular underpinnings of this process, we performed single-cell RNA-seq analysis of ADP-Hep-treated organoids as well as lineage-tracing experiments. Our data reveal that ADP-Hep induced the specific ablation of the homeostatic intestinal stem cell (ISC) pool. Removal of ADP-Hep resulted in the rapid recovery of ISCs through dedifferentiation of Paneth cells, which transiently acquired revSC features and expressed nuclear YAP. Moreover, lineage tracing from Lyz1+ Paneth cells showed that ADP-Hep triggered Paneth cell de-differentiation towards pluripotent and proliferative cells in organoids. In vivo, revSC emergence in response to irradiation-induced injury was severely blunted in Tifa-deficient mice, suggesting that efficient epithelial regeneration in this model required detection of microbiota-derived ADP-Hep by the ALPK1-TIFA pathway. Together, our work reveals that Paneth cells can serve as the cell of origin for revSC induction in the physiological context of microbial stimulation, and that the transient loss of Alpk1-expressing ISCs is the initiating event for this regenerative process.

cell biology↗

Restoring Protein Glycosylation with GlycoShape

During the past few years, we have been witnessing a revolution in structural biology. Leveraging on technological and computational advances, scientists can now resolve biomolecular structures at the atomistic level of detail by cryogenic electron microscopy (cryo-EM) and predict 3D structures from sequence alone by machine learning (ML). One technique often supports the other to provide the view of atoms in molecules required to capture the function of molecular machines. An example of the extraordinary impact of these advances on scientific discovery and on public health is given by how structural information supported the rapid development of COVID-19 vaccines based on the SARS-CoV-2 spike (S) glycoprotein. Yet, none of these new technologies can capture the details of the dense coat of glycans covering S, which is responsible for its natural, biologically active structure and function and ultimately for viral evasion. Indeed, glycosylation, the most abundant post-translational modification of proteins, is largely invisible through experimental structural biology and in turn it cannot be reproduced by ML, because of the lack of data to learn from. Molecular simulations through high-performance computing (HPC) can fill this crucial information gap, yet the computational resources, the users skills and the long timescales involved limit applications of molecular modelling to single study cases. To broaden access to structural information on glycans, here we introduce GlycoShape (https://glycoshape.org) an open access (OA) glycan structure database and toolbox designed to restore glycoproteins to their native functional form by supplementing the structural information available on proteins in public repositories, such as the RCSB PDB (www.rcsb.org) and AlphaFold Protein Structure Database (https://alphafold.ebi.ac.uk/), with the missing glycans derived from over 1 ms of cumulative sampling from molecular dynamics (MD) simulations. The GlycoShape Glycan Database (GDB) currently counts over 435 unique glycans principally covering the human glycome and with additional structures, fragments, and epitopes from other eukaryotic and prokaryotic organisms. The GDB feeds into Re-Glyco, a bespoke algorithm in GlycoShape designed to rapidly restore the natural glycosylation to protein 3D structures and to predict N-glycosylation occupancy, where unknown. Ultimately, integration of GlycoShape with other OA protein structure databases can provide a step-change in scientific discovery, from the structural and functional characterization of the active form of biomolecules, all the way down to pharmacological applications and drug discovery.

biophysics↗