Search bioRxiv⌕ Search

Biology subjects

Bonidia, R.

Publications and source records attributed to Bonidia, R..

2 recordsLinked to original sources

ANIMA: predicting protein-protein interactions across species

Motivation: Protein--protein interactions (PPIs) underpin a wide range of biological functions in living organisms. Experimental identification of new PPIs is expensive and time-consuming. The experimental bottleneck has implied an imbalance in terms of data availability: while certain species have been screened exhaustively, other species have not been sufficiently examined. An AI driven protocol for PPI prediction that leverages the massive data accumulated for certain species means a decisive boost for so far understudied species. Results: We present ANIMA (Artificial Neural Interaction Model for Animals), an AI supported cross-species PPI prediction model trained on popular species to predict PPIs in under-researched species. Our experiments demonstrate that our model, when trained on 200 diversely selected animal species, can successfully predict PPIs in other species: ANIMA achieves 95.3% accuracy on other animal species, 91.1% on other eukaryotes, and 82.5% on non-eukaryotes. For a more fine-grained evaluation of the model, we stratify performance rates by the evolutionary distance of test to training sets. We also stratify results by a novel, alignment based score ("representation score") which allows for fine-grained evaluation in terms of its capacity to generalize to unseen interactions. As expected, results demonstrate increasing performance on increasing evolutionary similarity and on increasing identity of interacting proteins, while still showing excellent performance on proteins entirely lacking counterparts in the training set. In comparison with the state of the art, ANIMA demonstrates substantial superiority in terms of performance rates.

bioinformatics↗

BioAutoML-FAST: an automated machine-learning platform for reusable and benchmarked biological sequence models

The prediction of biological sequence properties has traditionally relied on alignment-based methods that assume evolutionary homology and depend on curated reference databases. This, in turn, limits scalability and sensitivity for large or heterogeneous datasets, remote homologs, short sequences, and rapidly evolving genomic regions. Although Machine-Learning (ML) approaches offer alignment-free alternatives, their broader adoption is limited by: (i) the lack of standardized, externally validated benchmark models across diverse datasets, and (ii) the technical expertise required for feature engineering, model selection, and evaluation. Automated machine learning (AutoML) alleviates these challenges by systematically optimizing representations and models with minimal user intervention. However, most existing frameworks prioritize task-specific model construction and lack mechanisms for preserving trained models as persistent, comparable benchmarks. We introduce BioAutoML-FAST, an end-to-end web platform for automated ML analysis of nucleotide and amino acid sequences. It supports both classification and regression tasks and automates feature extraction, model training, and evaluation without requiring prior user expertise. Uniquely, it serves as a community benchmarking resource, hosting a continuously expanding repository of reusable, standardized models (currently 60) for genomic, transcriptomic, and proteomic applications. Extensive validation on independent datasets demonstrates performance comparable to or exceeding that of state-of-the-art methods, including protein language models such as ESM-2. BioAutoMLFAST is available at https://bioautoml.icmc.usp.br/. This website is free and open to all users, and there is no login requirement.

bioinformatics↗