Search bioRxiv⌕ Search

Biology subjects

Adduri, A.

Publications and source records attributed to Adduri, A..

3 recordsLinked to original sources

Predicting cellular responses to perturbation across diverse contexts with STATE

Cellular responses to perturbations are a cornerstone for understanding biological mechanisms and selecting drug targets. While machine learning models offer tremendous potential for predicting perturbation effects, they currently struggle to generalize to unobserved cellular contexts. Here, we introduce SO_SCPLOWTATEC_SCPLOW, a transformer model that predicts perturbation effects while accounting for cellular heterogeneity within and across experiments. SO_SCPLOWTATEC_SCPLOW predicts perturbation effects across sets of cells and is trained using gene expression data from over 100 million perturbed cells. SO_SCPLOWTATEC_SCPLOW improved discrimination of effects on large datasets by more than 30% and identified differentially expressed genes across genetic, signaling and chemical perturbations with significantly improved accuracy. Using its cell embedding trained on observational data from 167 million cells, SO_SCPLOWTATEC_SCPLOW identified strong perturbations in novel cellular contexts where no perturbations were observed during training. We further introduce Cell-Eval, a comprehensive evaluation framework that highlights SO_SCPLOWTATEC_SCPLOWs ability to detect cell type-specific perturbation responses, such as cell survival. Overall, the performance and flexibility of SO_SCPLOWTATEC_SCPLOW sets the stage for scaling the development of virtual cell models.

systems biology↗

Interpretable adenylation domain specificity prediction using protein language models

Natural products have long been a rich source of diverse and clinically effective drug candidates. Non-ribosomal peptides (NRPs), polyketides (PKs), and NRP-PK hybrids are three classes of natural products that display a broad range of bioactivities, including antibiotic, antifungal, anticancer, and immunosuppressant activities. However, discovering these compounds through traditional bioactivity-guided techniques is costly and time-consuming, often resulting in the rediscovery of known molecules. Consequently, genome mining has emerged as a high-throughput strategy to screen hundreds of thousands of microbial genomes to identify their potential to produce novel natural products. Adenylation domains play a key role in the biosynthesis of NRPs and NRP-PKs by recruiting substrates to incrementally build the final structure. We propose MASPR, a machine learning method that leverages protein language models for accurate and interpretable predictions of A-domain substrate specificities. MASPR demonstrates superior accuracy and generalization over existing methods and is capable of predicting substrates not present in its training data, or zero-shot classification. We use MASPR to develop Seq2Hybrid, an efficient algorithm to predict the structure of hybrid NRP-PK natural products from microbial genomes. Using Seq2Hybrid, we propose putative biosynthetic gene clusters for the orphan natural products Octaminomycin A, Dityromycin, SW-163B, and JBIR-39.

bioinformatics↗

Ornaments for Accurate and Efficient Allele-Specific Expression Estimation with Bias Correction

Allele-specific expression has been used to elucidate various biological mechanisms, such as genomic imprinting and gene expression variation caused by genetic changes in cis-regulatory elements. However, existing methods for obtaining allele-specific expression from RNA-seq reads do not adequately and efficiently remove various biases, such as reference bias, where reads containing the alternative allele do not map to the reference transcriptome, or ambiguous mapping bias, where reads containing the reference allele map differently from reads containing the alternative allele. We present Ornaments, a computational tool for rapid and accurate estimation of allele-specific expression at unphased heterozygous loci from RNA-seq reads while correcting for allele-specific read mapping bias. Ornaments removes reference bias by accounting for personalized transcriptome, and ambiguous mapping bias by probabilistically assigning reads to multiple transcripts and variant loci they map to. Ornaments is a lightweight extension of kallisto, a popular tool for fast RNA-seq quantification, that improves the efficiency and accuracy of WASP, a popular tool for bias correction in allele-specific read mapping. Our experiments on simulated and human lymphoblastoid cell-line RNA-seq reads with the genomes of the 1000 Genomes Project show that Ornaments is as efficient as kallisto, an order of magnitude faster than WASP, and more accurate than WASP and kallisto. In addition, Ornaments detected genes that are imprinted at transcript level with higher sensitivity, compared to WASP that detected the imprinted signals only at gene level.

genetics↗