Search bioRxiv⌕ Search

Biology subjects

Haas, N.

Publications and source records attributed to Haas, N..

2 recordsLinked to original sources

TransStop, a genomic language model for the pan-drug prediction of translational readthrough efficacy

MotivationPremature termination codons (PTCs) are a major cause of genetic diseases, but the efficacy of therapeutic readthrough agents is highly context-dependent. While linear models have shown promise in predicting readthrough efficiency, they may not fully capture the complex, non-linear interactions between sequence context and drug activity. MethodsWe developed TransStop, a transformer-based pan-drug model, trained on a dataset of [~]5,400 PTCs and eight readthrough compounds. The model learns sequence representations and incorporates learnable embeddings to capture drug-specific effects, allowing a single model to predict efficacy for multiple drugs. ResultsOur model achieved a global R2=0.94 on a held-out test set. Visualizations of the learned embeddings revealed a deep understanding of biological principles, including the distinct clustering of stop codon types and grouping of drugs by mechanism of action. In silico saturation mutagenesis and epistasis analyses uncovered complex, non-additive sequence determinants of readthrough. We generated 32.7 million predictions across the human genome, covering all possible PTCs. Analysis of these genome-wide predictions revealed strong drug specializations for specific stop codon contexts and identified key areas of disagreement with previous models, particularly for UGA codons, where our model predicts a more effective drug in thousands of cases. The TransStop model represents a significant advancement in the prediction of translational readthrough efficiency. Its superior accuracy and the biological insights derived from its applications provide a powerful tool for guiding clinical trial design, drug development, and personalized patient treatment. Availability and ImplementationSource code: https://github.com/Dichopsis/TransStop. Model: https://huggingface.co/Dichopsis/TransStop.Genome-wide predictions: https://doi.org/10.5281/zenodo.16918476.

bioinformatics↗

Crowdsourced Protein Design: Lessons From the Adaptyv EGFR Binder Competition

In this report, we summarize and analyze the 2024 Adaptyv protein design competition. Participants used computational and Machine Learning (ML) methods of their choice to design proteins that bind the Epidermal Growth Factor Receptor (EGFR), a key drug target involved in cell growth, differentiation, and cancer development. Over 1,800 designs were submitted across two rounds. Of these, 601 proteins were selected and characterized for expression and binding affinity to EGFR, with competitors both optimizing existing binders (KD = 1.21 nM) and creating de novo binders (KD = 82 nM). All selected designs were experimentally validated using Adaptyvs automated Bio-Layer Interferometry (BLI) pipeline. This competition illustrates the potential of crowdsourcing to drive creativity and innovation in protein design. However, it also exposed key challenges, such as the lack of standardized benchmarks, experimental design targets, and robust computational metrics for method comparison. We anticipate that future competitions will address these gaps and further motivate progress in computational protein design.

bioengineering↗