Search bioRxiv⌕ Search

Biology subjects

Ahuja, G.

Publications and source records attributed to Ahuja, G..

5 recordsLinked to original sources

A visual atlas of genes tissue-specific pathological roles

Dysregulation of a genes function, either due to mutations or impairments in regulatory networks, often triggers pathological states in the affected tissue. Comprehensive mapping of these apparent gene-pathology relationships is an ever daunting task, primarily due to genetic pleiotropy and lack of suitable computational approaches. With the advent of high throughput genomics platforms and community scale initiatives such as the Human Cell Landscape (HCL) project [1], researchers have been able to create gene expression portraits of healthy tissues resolved at the level of single cells. However, a similar wealth of knowledge is currently not at our finger-tip when it comes to diseases. This is because the genetic manifestation of a disease is often quite heterogeneous and is confounded by several clinical and demographic covariates. To circumvent this, we mined ~18 million PubMed abstracts published till May 2019 and selected ~6.1 million of them that describe the pathological role of genes in different diseases. Further, we employed a word embedding technique from the domain of Natural Language Processing (NLP) to learn vector representation of entities such as genes, diseases, tissues, etc., in a way such that their relationship is preserved in a vector space. Notably, Pathomap, by the virtue of its underpinning theory, also learns transitive relationships. Pathomap provided a vector representation of words indicating a possible association between DNMT3A/BCOR with CYLD cutaneous syndrome (CCS). The first manuscript reporting this finding was not part of our training data. Key pointsO_LIWe mined ~18 million PubMed abstracts to extract latent knowledge pertaining to tissue specific pathological roles of genes. C_LIO_LIWe found well-defined gene modules implicated in disease pathogenesis in anatomically proximal tissues. C_LIO_LIWe demonstrated an ahead of time discovery of the association between DNMT3A/BCOR with CYLD cutaneous syndrome (CCS), as a knowledge synthesis use-case. C_LI

bioinformatics↗

Artificial Intelligence uncovers carcinogenic human metabolites

The genome of a eukaryotic cell is often vulnerable to both intrinsic and extrinsic threats due to its constant exposure to a myriad of heterogeneous compounds. Despite the availability of innate DNA damage response pathways, some genomic lesions trigger cells for malignant transformation. Accurate prediction of carcinogens is an ever-challenging task due to the limited information about bona fide (non)carcinogens. We developed Metabokiller, an ensemble classifier that accurately recognizes carcinogens by quantitatively assessing their electrophilicity as well as their potential to induce proliferation, oxidative stress, genomic instability, alterations in the epigenome, and anti-apoptotic response. Concomitant with the carcinogenicity prediction, Metabokiller is fully interpretable since it reveals the contribution of the aforementioned biochemical properties in imparting carcinogenicity. Metabokiller outperforms existing best-practice methods for carcinogenicity prediction. We used Metabokiller to unravel cells endogenous metabolic threats by screening a large pool of human metabolites and predicted a subset of these metabolites that could potentially trigger malignancy in normal cells. To cross-validate Metabokiller predictions, we performed a range of functional assays using Saccharomyces cerevisiae and human cells with two Metabokiller-flagged human metabolites namely 4-Nitrocatechol and 3,4-Dihydroxyphenylacetic acid and observed high synergy between Metabokiller predictions and experimental validations.

bioinformatics↗

Marker-free characterization of single live circulating tumor cell full-length transcriptomes

The identification and characterization of circulating tumor cells (CTCs) are important for gaining insights into the biology of metastatic cancers, monitoring disease progression, and medical management of the disease. The limiting factor that hinders enrichment of purified CTC populations is their sparse availability, heterogeneity, and altered phenotypic traits relative to the tumor of origin. Intensive research both at the technical and molecular fronts led to the development of assays that ease CTC detection and identification from the peripheral blood. Most CTC detection methods use a mix of size selection, immune marker based white blood cells (WBC) depletion, and positive enrichment antibodies targeting tumor-associated antigens. However, the majority of these methods either miss out on atypical CTCs or suffer from WBC contamination. Single-cell RNA sequencing (scRNA-Seq) of CTCs provides a wealth of information about their tumors of origin as well as their fate and is a potent method of enabling unbiased identification of CTCs. We present unCTC, an R package for unbiased identification and characterization of CTCs from single-cell transcriptomic data. unCTC features many standard and novel computational and statistical modules for various analysis tasks. These include a novel method of scRNA-Seq clustering, named Deep Dictionary Learning using K-means clustering cost (DDLK), expression based copy number variation (CNV) inference, and combinatorial, marker-based verification of the malignant phenotypes. DDLK enables robust segregation of CTCs and WBCs in the pathway space, as opposed to the gene expression space. We validated the utility of unCTC on scRNA-Seq profiles of breast CTCs from six patients, captured and profiled using an integrated ClearCell(R) FX and PolarisTM workflow that works by the principles of size-based separation of CTCs and marker based WBC depletion.

bioinformatics↗

Gene expression based inference of drug resistance in cancer

Inter and intra-tumoral heterogeneity are major stumbling blocks in the treatment of cancer and are responsible for imparting differential drug responses in cancer patients. Recently, the availability of large-scale drug screening datasets has provided an opportunity for predicting appropriate patient-tailored therapies by employing machine learning approaches. In this study, we report a predictive modeling approach to infer treatment response in cancers using gene expression data. In particular, we demonstrate the benefits of considering integrated chemogenomics approach, utilizing the molecular drug descriptors and pathway activity information as opposed to gene expression levels. We performed extensive validation of our approach on tissue-derived single-cell and bulk expression data. Further, we constructed several prostate cancer cell lines and xenografts, exposed to differential treatment conditions to assess the predictability of the outcomes. Our approach was further assessed on pan-cancer RNA-sequencing data from The Cancer Genome Atlas (TCGA) archives, as well as an independent clinical trial study describing the treatment journey of three melanoma patients. To summarise, we benchmarked the proposed approach on cancer RNA-seq data, obtained from cell lines, xenografts, as well as humans. We concluded that pathway-activity patterns in cancer cells are reasonably indicative of drug resistance, and therefore can be leveraged in personalized treatment recommendations.

cancer biology↗

Translational specialization in pluripotency by RBPMS poises future lineage-decisions

The blueprints for developing organs are preset at the early stages of embryogenesis. Transcriptional and epigenetic mechanisms are proposed to preset developmental trajectories. However, we reveal that the competence for future cardiac fate of human embryonic stem cells (hESCs) is preset in pluripotency by a specialized mRNA translation circuit controlled by RBPMS. RBPMS is recruited to active ribosomes in hESCs to control the translation of essential factors needed for cardiac commitment program, including WNT signaling. Consequently, RBPMS loss specifically and severely impedes cardiac mesoderm specification leading to patterning and morphogenesis defects in human cardiac organoids. Mechanistically, RBPMS specializes mRNA translation, selectively via 3UTR binding and globally by promoting translation initiation. Accordingly, RBPMS loss causes translation initiation defects highlighted by aberrant retention of the EIF3 complex and depletion of EIF5A from mRNAs, thereby abrogating ribosome recruitment. We reveal how future fate trajectories are preprogrammed during embryogenesis by specialized mRNA translation. Teaser: Cardiac fate competence is preprogrammed in pluripotency by specialized mRNA translation of factors initiating cardiogenesis

developmental biology↗