Search bioRxiv⌕ Search

Biology subjects

Fazel-Zarandi, M.

Publications and source records attributed to Fazel-Zarandi, M..

2 recordsLinked to original sources

Integrative Computational Framework, Dyscovr, Links Mutated Driver Genes to Expression Dysregulation Across 19 Cancer Types

Mutations within cancer driver genes induce widespread transcriptional changes that reflect altered cellular states and can reveal associated genetic vulnerabilities. However, it remains challenging to determine which genes are dysregulated as a consequence of cancer alterations, and of these, which represent therapeutic opportunities. Here, we present Dyscovr, an integrative computational framework that leverages somatic mutation, gene expression, copy number alteration, methylation, and clinical data from primary tumors to identify driver-associated transcriptional changes. Dyscovr then uses these transcriptional changes as a biologically grounded starting point, integrating them with cancer cell line data to prioritize genes whose inhibition is predicted to reduce viability either specifically in driver-mutant contexts or in combination with driver inhibition. Applied both pan-cancer and across 19 tumor types, Dyscovr uncovers hundreds of such conditional vulnerabilities. As a case study, we newly implicate--and experimentally validate--KBTBD2 as a gene whose inhibition enhances the efficacy of PI3K inhibitors in PIK3CA-mutant breast cancer cell lines. The Dyscovr software (github.com/Singh-Lab/Dyscovr) and predictions (dyscovr.princeton.edu) provide a platform and resource for linking mutated driver genes to conditional genetic vulnerabilities and for prioritizing these relationships for experimental and therapeutic investigation.

systems biology↗

Language models of protein sequences at the scale of evolution enable accurate structure prediction

Artificial intelligence has the potential to open insight into the structure of proteins at the scale of evolution. It has only recently been possible to extend protein structure prediction to two hundred million cataloged proteins. Characterizing the structures of the exponentially growing billions of protein sequences revealed by large scale gene sequencing experiments would necessitate a break-through in the speed of folding. Here we show that direct inference of structure from primary sequence using a large language model enables an order of magnitude speed-up in high resolution structure prediction. Leveraging the insight that language models learn evolutionary patterns across millions of sequences, we train models up to 15B parameters, the largest language model of proteins to date. As the language models are scaled they learn information that enables prediction of the three-dimensional structure of a protein at the resolution of individual atoms. This results in prediction that is up to 60x faster than state-of-the-art while maintaining resolution and accuracy. Building on this, we present the ESM Metage-nomic Atlas. This is the first large-scale structural characterization of metagenomic proteins, with more than 617 million structures. The atlas reveals more than 225 million high confidence predictions, including millions whose structures are novel in comparison with experimentally determined structures, giving an unprecedented view into the vast breadth and diversity of the structures of some of the least understood proteins on earth.

synthetic biology↗