Search bioRxiv⌕ Search

Biology subjects

Challacombe, C. A.

Publications and source records attributed to Challacombe, C. A..

2 recordsLinked to original sources

Crowdsourced Protein Design: Lessons From the Adaptyv EGFR Binder Competition

In this report, we summarize and analyze the 2024 Adaptyv protein design competition. Participants used computational and Machine Learning (ML) methods of their choice to design proteins that bind the Epidermal Growth Factor Receptor (EGFR), a key drug target involved in cell growth, differentiation, and cancer development. Over 1,800 designs were submitted across two rounds. Of these, 601 proteins were selected and characterized for expression and binding affinity to EGFR, with competitors both optimizing existing binders (KD = 1.21 nM) and creating de novo binders (KD = 82 nM). All selected designs were experimentally validated using Adaptyvs automated Bio-Layer Interferometry (BLI) pipeline. This competition illustrates the potential of crowdsourcing to drive creativity and innovation in protein design. However, it also exposed key challenges, such as the lack of standardized benchmarks, experimental design targets, and robust computational metrics for method comparison. We anticipate that future competitions will address these gaps and further motivate progress in computational protein design.

bioengineering↗

Towards a Dataset for State of the Art Protein Toxin Classification

In-silico toxin classification assists in industry and academic endeavors and is critical for biosecurity. For instance, proteins and peptides hold promise as therapeutics for a myriad of conditions, and screening these biomolecules for toxicity is a necessary component of synthesis. Additionally, with the expanding scope of biological design tools, improved toxin classification is essential for mitigating dual-use risks. Here, a general toxin classifier that is capable of addressing these demands is developed. Applications for in-silico toxin classification are discussed, conventional and contemporary methods are reviewed, and criteria defining current needs for general toxin classification are introduced. As contemporary methods and their datasets only partially satisfy these criteria, a comprehensive approach to toxin classification is proposed that consists of training and validating a single sequence classifier, BioLMTox, on an improved dataset that unifies current datasets to align with the criteria. The resulting benchmark dataset eliminates ambiguously labeled sequences and allows for direct comparison against nine previous methods. Using this comprehensive dataset, a simple fine-tuning approach with ESM-2 was employed to train BioLMTox, resulting in accuracy and recall validation metrics of 0.964 and 0.984, respectively. This LLM-based model does not use traditional alignment methods and is capable of identifying toxins of various sequence lengths from multiple domains of life in sub-second time frames.

synthetic biology↗