Search bioRxiv⌕ Search

Biology subjects

Reade, W.

Publications and source records attributed to Reade, W..

3 recordsLinked to original sources

Advances in protein function prediction from the fifth CAFA challenge

The Critical Assessment of Functional Annotation (CAFA) is a long-standing community effort to independently assess computational methods for protein function prediction, to highlight well-performing methodologies, to identify bottlenecks in the field, and to provide a forum for the dissemination of results and exchange of ideas. In its fifth round (CAFA5) of triennial challenges, a partnership with Kaggle Inc. facilitated participation from a large community of data scientists and computational biologists through a competitive prospective challenge on the crowdsourcing platform. In this work, we present an in-depth analysis of the submitted predictions and report improvements in accuracy over all methods from the previous CAFA challenges. We further introduce a new evaluation setting for proteins with pre-existing (incomplete) annotations and identify the need for methods that better leverage existing annotations to predict those that will be discovered later. Finally, we characterize the prospective evaluation framework by examining performance on a strict set of unpublished annotations and across intermediate database releases. Our results indicate that recent developments in the field, such as the availability of protein language models and accurately predicted 3D structures, as well as the growth of experimental annotations through biocuration, have all contributed to performance improvements. 1

bioinformatics↗

Lessons learned from a Kaggle challenge for particle picking in cryo-electron tomography

The difficulty of particle picking in cryo-electron tomography remains a barrier to routine in situ structure determination. Machine learning is well-suited to overcome this bottleneck with efficient algorithms that generalize across molecular species. To spur new algorithm development, we held a three-month Kaggle challenge that tasked contestants with annotating five molecular species across hundreds of experimental tomograms. This competition successfully engaged >1000 participants from diverse fields and delivered particle pickers that outperformed existing state-of-the-art. Systematic comparisons of the contestants submissions revealed the tolerance of subtomogram averaging to moderate but not severe over-picking and underscored the need for more robust measures of annotation quality. The winning models also highlighted the importance of data augmentation to overcome limited training data. All competition tomograms along with the ground truth and winning teams annotations have been released on the CryoET Data Portal as a resource to benchmark current and future tools for particle picking.

cell biology↗

Ribonanza: deep learning of RNA structure through dual crowdsourcing

Prediction of RNA structure from sequence remains an unsolved problem, and progress has been slowed by a paucity of experimental data. Here, we present Ribonanza, a dataset of chemical mapping measurements on two million diverse RNA sequences collected through Eterna and other crowdsourced initiatives. Ribonanza measurements enabled solicitation, training, and prospective evaluation of diverse deep neural networks through a Kaggle challenge, followed by distillation into a single, self-contained model called RibonanzaNet. When fine tuned on auxiliary datasets, RibonanzaNet achieves state-of-the-art performance in modeling experimental sequence dropout, RNA hydrolytic degradation, and RNA secondary structure, with implications for modeling RNA tertiary structure.

biophysics↗