Search bioRxiv⌕ Search

Biology subjects

Gonzalez-Puelma, J.

Publications and source records attributed to Gonzalez-Puelma, J..

2 recordsLinked to original sources

Data-Centric Evaluation of Protein Function Prediction Pipelines

Performance estimates in protein function prediction depend not only on model choice but also on upstream decisions that define the learning problem. Using antioxidant protein classification as a controlled case study, we evaluated how dataset harmonisation, protein representation, redundancy control, and partitioning strategy affect protein machine learning pipelines. We integrated 18,804 records from 12 publicly available dataset entries into a curated consensus dataset of 4,193 protein sequences. One-hot encoding and six pretrained protein language model representations were evaluated as model inputs and as similarity spaces for redundancy reduction and distance-aware splitting. Representation choice substantially altered dataset geometry, retained dataset size, class balance, and downstream evaluation. At representation-specific p90 thresholds, one-hot encoding retained the complete dataset, whereas pretrained embeddings retained between 5% and 25% of sequences. Distance-aware partitioning reduced apparent performance relative to random splitting by up to 0.15 MCC before redundancy control, while this difference narrowed after similarity filtering. Selected configurations nevertheless maintained high performance under stricter evaluation, reaching an MCC of 0.84. These findings show that performance estimates should be interpreted as outcomes of complete data-centric workflows rather than isolated properties of predictive models.

bioinformatics↗

MAOMAO: An Ontology-Guided FAIR Resource for Harmonized Peptide Toxicity Data

Peptide toxicity is a critical safety and developability parameter in peptide discovery and therapeutic development, yet relevant information remains fragmented across databases, literature resources, and curated datasets. Here, we present MAOMAO, an ontology-guided FAIR-oriented resource that integrates and harmonizes peptide toxicity data from 54 sources. MAOMAO contains 71,701 unique peptide sequences across seven toxicity-related endpoints, represented as 501,907 sequence endpoint combinations with endpoint-specific evidence states and explicit encoding of unavailable information. The resource combines standardized terminology, a hierarchical toxicity vocabulary, evidence-aware state resolution, provenance-aware metadata, 41 physicochemical descriptors, 10 protein language model representations, and a one-hot baseline. It provides endpoint-specific benchmark partitions across splitting strategies and random seeds, reusable with numerical representations. MAOMAO establishes a reusable framework for peptide toxicity research and data-driven toxicology.

pharmacology and toxicology↗