Search bioRxiv⌕ Search

Biology subjects

Coley, C. W.

Publications and source records attributed to Coley, C. W..

6 recordsLinked to original sources

Identification of Antituberculars with Favorable Potency and Pharmacokinetics through Structure-Based and Ligand-Based Modeling

Drug discovery is inherently challenged by a multiple criteria decision making problem. The arduous path from hit discovery through lead optimization and preclinical candidate selection necessitates the evolution of a plethora of molecular properties. In this study, we focus on the hit discovery phase while beginning to address multiple criteria critical to the development of novel therapeutics to treat Mycobacterium tuberculosis infection. We develop a hybrid structure- and ligand-based pipeline for nominating diverse inhibitors targeting the {beta}-ketoacyl synthase KasA by employing a Bayesian optimization-guided docking method and an ensemble model for compound nominations based on machine learning models for in vitro antibacterial efficacy, as characterized by minimum inhibitory concentration (MIC), and mouse pharmacokinetic (PK) plasma exposure. The application of our pipeline to the Enamine HTS library of 2.1M molecules resulted in the selection of 93 compounds, the experimental validation of which revealed exceptional PK (41%) and MIC (19%) success rates. Twelve compounds meet hit-like criteria in terms of MIC and PK profile and represent promising seeds for future drug discovery programs.

microbiology↗

EnzymeCAGE: A Geometric Foundation Model for Enzyme Retrieval with Evolutionary Insights

Enzyme catalysis is fundamental to life, driving the chemical transformations that sustain biological processes and support industrial applications. However, unraveling the intertwined relationships between enzymes and their catalytic reactions remains a significant challenge. Here, we present EnzymeCAGE, a catalytic-specific geometric foundation model trained on approximately 1 million structure-informed enzyme-reaction pairs, spanning over 2,000 species and encompassing an extensive diversity of genomic and metabolic information. EnzymeCAGE features a geometry-aware multi-modal architecture coupled with an evolutionary information integration module, enabling it to effectively model the nuanced relationships between enzyme structure, catalytic function, and reaction specificity. EnzymeCAGE supports both experimental and predicted enzyme structures and is applicable across diverse enzyme families, accommodating a broad range of metabolites and reaction types. Extensive evaluations demonstrate EnzymeCAGEs state-of-the-art performance in enzyme function prediction, reaction de-orphaning, catalytic site identification, and biosynthetic pathway reconstruction. These results highlight its potential as a transformative foundation model for understanding enzyme catalysis and accelerating the discovery of novel biocatalysts.

synthetic biology↗

ema-tool: a Python Library for the Comparative Analysis of Embeddings from Biomedical Foundation Models

Foundation models, which encode patterns in large, high-dimensional data as embeddings, show promise in many machine learning related applications in molecular biology. Embeddings learned by the models provide informative features for downstream prediction tasks, however, the information captured by the model is often not interpretable. One approach to understanding the captured information is through the analysis of their learned embeddings, which in molecular biology so far has mainly focused on visualizing individual embedding spaces. This study introduces a quantitative framework for cross-space comparison, enabling intuitive exploration and comparison of embedding spaces in molecular biology. The framework emphasizes analyzing the distribution of known biological information within embedding space neighborhoods and provides insights into relationships between multiple embedding spaces. Comparison techniques include global pairwise distance measurements as well as local nearest neighbor analyses. By applying our framework to embeddings from protein language models, we demonstrate how embedding space analysis can serve as a valuable pre-filtering step for task-specific supervised machine learning applications and for the recognition of differential patterns in data encoded within and across different embedding spaces. To support a wide usability, we provide a Python library that implements all analysis methods, available at https://github.com/broadinstitute/EmmaEmb.

bioinformatics↗

Annotating metabolite mass spectra with domain-inspired chemical formula transformers

Metabolomic studies have succeeded in identifying small molecule metabolites that mediate cell signaling, competition, and disease pathology in part due to large-scale community efforts to measure mass spectra for thousands of metabolite standards. Nevertheless, the vast majority of spectra observed in clinical samples cannot be unambiguously matched to known structures, suggesting powerful opportunities for further discoveries in the dark metabolome. Deep learning approaches to small molecule structure elucidation have surprisingly failed to rival classical statistical methods, which we hypothesize is due to the lack of in-domain knowledge incorporated into current neural network architectures. We introduce a new neural network driven workflow for untargeted metabolomics, Metabolite Inference with Spectrum Transformers (MIST), to annotate mass spectrometry peaks with chemical structures generalizing beyond known standards. Unlike other neural approaches, MIST incorporates domain insights into its architecture by forcing the network to more directly link peaks to physical atom representations, neutral losses, and chemical substructures. MIST outperforms both standard neural architectures and the state-of-the-art kernel method on fingerprint prediction from spectra for over 70% of metabolite standards and retrieves over 66% of metabolites with equal or improved accuracy, with 29% strictly better. We further demonstrate the utility of MIST in a prospective setting to identify new differentially abundant metabolite structures from an inflammatory bowel disease patient cohort and subsequently annotate dipeptides and alkaloid compounds without spectral standards.

bioinformatics↗

DNA-encoded library (DEL)-enabled discovery of proximity-inducing small molecules

Molecular glues and bifunctional compounds that induce protein-protein associations provide a powerful and general mechanism to modulate cell circuitry. We sought to develop a platform for the direct discovery of compounds able to induce association of any two pre-selected proteins, using the first bromodomain of BRD4 and the VHL-elongin C-elongin B (VCB) complex as a test system. Leveraging the screening power of DNA-encoded libraries (DELs), we synthesized [~]one million DNA-encoded compounds that possess a VHL-targeting fragment, a variety of connectors, and a diversity element generated by split- and-pool combinatorial chemistry. By screening our DEL against BRD4BD1 in the presence and absence of VCB, we could identify VHL-bound molecules that simultaneously bind BRD4. For highly barcode-enriched library members, ternary complex formation leading to BRD4 degradation was confirmed in cells. Furthermore, a ternary complex crystal structure was obtained for the most enriched library member. Our work provides a foundation for adapting DEL screening to the discovery of proximity-inducing small molecules. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=73 SRC="FIGDIR/small/512184v1_ufig1.gif" ALT="Figure 1"> View larger version (22K): org.highwire.dtl.DTLVardef@43f3c9org.highwire.dtl.DTLVardef@13a1f9eorg.highwire.dtl.DTLVardef@f1d751org.highwire.dtl.DTLVardef@16f3789_HPS_FORMAT_FIGEXP M_FIG C_FIG

pharmacology and toxicology↗

Diversity-oriented synthesis encoded by deoxyoligonucleotides

Diversity-oriented synthesis (DOS)is a powerful strategy to prepare molecules with underrepresented features in commercial screening collections, resulting in the elucidation of novel biological mechanisms. In parallel to the development of DOS, DNA-encoded libraries (DELs) have emerged as an effective, efficient screening strategy to identify protein binders. Despite recent advancements in this field, most DEL syntheses are limited by the presence of sensitive DNA-based constructs. Here, we describe the design, synthesis, and validation experiments performed for a 3.7 million-member DEL, generated using diverse skeleton architectures with varying exit vectors, derived from DOS, to achieve structural diversity beyond what is possible by varying appendages alone. We will make this DEL available to the academic scientific community to increase access to novel structural features and accelerate early-phase drug discovery.

biochemistry↗