Search bioRxiv⌕ Search

Biology subjects

Afonso, M. Q. L.

Publications and source records attributed to Afonso, M. Q. L..

3 recordsLinked to original sources

The Human Bindome: A Proteome-scale Atlas of Designed Binder Candidates

Affinity reagents such as antibodies are indispensable for interrogating proteins biological function. Yet they are costly and frequently unreliable, with unknown sequences, posing challenges to reproducible experimental research. Deep learning-based protein design can now in silico generate affinity reagents achieving reliable experimental success rates, but has remained largely confined to specialist laboratories. Here we present the Human Bindome, a proteome-scale atlas of high-confidence in silico protein binder candidates. By embedding the experimentally benchmarked BindCraft method in an accelerated, parallelized framework with automated domain-level target selection, we generated 306,146 binder candidates covering 8,296 human proteins (40.9% of the full proteome). Every candidate carries a defined sequence, a predicted binder-target structure model, and in silico confidence metrics. We characterize proteome-wide coverage and show that binder epitopes frequently overlap functional sites. This positions the Bindome as a resource of genetically encodable perturbagens for site-specific, modular control of protein function. The Bindome is freely available through a web interface (https://bindome.epfl.ch), with agentic, natural-language querying and as data splits for machine-learning model development. We anticipate that the Bindome will be valuable for the scientific community by providing affinity and perturbation reagents with broad applications in dissecting biological mechanisms as well as in drug and target discovery.

synthetic biology↗

PDBe tools for an in-depth analysis of small molecules in the Protein Data Bank

The Protein Data Bank (PDB) is the primary global repository for experimentally determined 3D structures of biological macromolecules and their complexes with ligands, proteins, and nucleic acids. PDB contains over 47,000 unique small molecules bound to the macromolecules. Despite the extensive data available, the complexity of small molecule data in the PDB necessitates specialised tools for effective analysis and visualisation. PDBe has developed a number of tools, including PDBe CCDUtils (https://github.com/PDBeurope/ccdutils) for accessing and enriching ligand data, PDBe Arpeggio (https://github.com/PDBeurope/arpeggio) for analysing interactions between ligands and macromolecules, and PDBe RelLig (https://github.com/PDBeurope/rellig) for identifying the functional roles of ligands (such as reactants, cofactors, or drug-like molecules) within protein-ligand complexes. Furthermore, the enhanced ligand annotations and data generated by these tools are presented in a comprehensive view on the novel PDBe-KB ligand pages, providing a holistic view of small molecules that enables the establishment of their biological contexts (Example page for Imatinib: https://wwwdev.ebi.ac.uk/pdbe-srv/pdbechem/chemicalCompound/show/STI). By improving the standardisation of ligand identification, adding various annotations, and offering advanced visualisation capabilities, these tools help researchers navigate the complexities of small molecules and their roles in biological systems, facilitating mechanistic understanding of biological functions. The ongoing enhancements to these resources are designed to support the scientific community in gaining valuable insights into ligands and their applications across various fields, including drug discovery, molecular biology, systems biology, structural biology, and pharmacology.

bioinformatics↗

Human-in-the-loop approach to identify functionally important residues of proteins from literature

We present a novel system that leverages curators in the loop to develop a dataset and model for detecting residue-level functional annotations and other protein structure features from standard publication text. Our approach involves the integration of data from multiple resources, including PDBe, EuropePMC, PubMedCentral, and PubMed, combined with annotation guidelines from UniProt, while employing LitSuggest and Huggingface models as tools in the annotation process. A team of seven annotators manually curated ten articles for named entities, which we utilized to train a starting PubmedBert model from Huggingface. Using a human-in-the-loop annotation system, we developed the best model with commendable performance metrics of 0.90 for precision, 0.92 for recall, and 0.91 for F1-measure. Our proposed system showcases a successful synergy of machine learning techniques and human expertise in curating a dataset for residue-level functional annotations and protein structure features. The results demonstrate the potential for broader applications in protein research, bridging the gap between advanced machine learning models and the indispensable insights of domain experts.

bioinformatics↗