Search bioRxiv⌕ Search

Biology subjects

De Paolis Klauza, M. C.

Publications and source records attributed to De Paolis Klauza, M. C..

2 recordsLinked to original sources

On the state of protein function prediction: a report on the fourth CAFA challenge

BackgroundThe Critical Assessment of Functional Annotation (CAFA) is a community effort held to understand the field of computational protein function prediction. Every three years, since 2010, the organizers initiate an experiment to collect function predictions on a large set of proteins and then evaluate the performance of predicting methods on a subset of proteins that have accumulated experimental annotations between the submission deadline and the evaluation time. CAFA provides an independent and rigorous assessment of the current state of the art, thus leveling the playing field, highlighting successes, revealing bottlenecks, and offering a forum for the exchange of ideas in protein science. Here, we report the results of the fourth CAFA experiment (CAFA4). ResultsCAFA4 featured the participation of 148 methods from 70 research groups on a total of 46,205 unique proteins over a 5-year annotation accumulation phase, the longest in any CAFA. In a comparison across CAFA2-CAFA4 methods, the prediction of Gene Ontology (GO) terms has clearly improved across all three GO aspects and traditional evaluation settings. While not achieving the first rank, several CAFA2 and CAFA3 methods featured in the top ten methods in many evaluations, suggesting that earlier methods still hold relevance. The performance is weaker in the newly introduced "partial knowledge" evaluation category (proteins with experimental annotations before submission deadline that gained additional annotations in the same GO aspect during the annotation accumulation phase), highlighting the need for a new class of methods. The rankings of the methods were stable over the years in traditional evaluation settings, but less so in the new partial knowledge evaluation. Overall, the field continues to progress with some influx of new participants. Sustained efforts will be necessary to substantially advance it.

bioinformatics↗

16S rRNA k-mer composition encodes microbial functional potential

16S rRNA amplicon sequencing is widely used for microbiome profiling, but most methods rely on reference databases of characterized organisms, limiting its accuracy in function prediction for underrepresented environments. We discovered that 16S rRNA k-mer composition carries substantial functional signal: (i) whole-genome k-mer profiles predict genome-encoded functions, and (ii) 16S rRNA k-mer profiles reflect their source genomes composition. Building on these relationships, we developed embeRNA, a neural network framework that predicts functions directly from 16S rRNA k-mer embeddings without requiring taxonomy assignment or phylogenetic placement. embeRNA outputs per-function probability scores, enabling users to tune decision thresholds to balance precision and recall or account for community novelty. In a stringent "novel microbes" benchmark - where all test sequences shared <97% identity with training data - embeRNA outperformed reference-based methods, particularly for hard-to-label functions. Applied to soil metagenomes with paired 16S and whole metagenome shotgun sequencing (WMS) data, embeRNA recovered most WMS-inferred functions and produced abundance profiles strongly correlated with WMS results, attaining better performance than a reference-based approach. Our findings demonstrate that 16S rRNA directly captures functional potential, and 16S amplicon sequencing data can complement WMS-based inference to broaden functional characterization of microbiomes, especially in understudied environments.

bioinformatics↗