bioRxiv · 10.1101/2024.08.02.606297
Importance of updated benchmark sets for statistically correct AlphaFold applications
Abstract
AlphaFold2 changed structural biology by providing high-quality structure predictions for all possible proteins. Since its inception, a plethora of applications were built on AlphaFold2, expediting discoveries in virtually all areas related to protein science. In many cases, however, optimism seems to have made scientists forget about data leakage, a serious issue that needs to be addressed when evaluating machine learning methods. Here we provide a rigorous benchmark set that can be used in a broad range of applications built around AlphaFold2/3. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=87 SRC="FIGDIR/small/606297v3_ufig1.gif" ALT="Figure 1"> View larger version (18K): org.highwire.dtl.DTLVardef@145b4deorg.highwire.dtl.DTLVardef@1656df0org.highwire.dtl.DTLVardef@14a94dorg.highwire.dtl.DTLVardef@773bd2_HPS_FORMAT_FIGEXP M_FIG C_FIG Key PointsO_LIWhen building applications on AlphaFold, scientists should consider the possibility of data leakage between AlphaFold training set and the independent test set of their method C_LIO_LIBETA provides multiple datasets with structures and sequences that were not used during the training of AlphaFold C_LIO_LIThese datasets provide a diverse range of use cases C_LIO_LIThe protocol was applied when building a simple disordered prediction method, showing different parameters required to optimize disordered prediction for proteins not used in AlphaFold training C_LI
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Dobson, L., Tusnady, G. E., Tompa, P.. 2024-08-06. Importance of updated benchmark sets for statistically correct AlphaFold applications. https://doi.org/10.1101/2024.08.02.606297
Cite the original work for its findings. Save a collection to share your selection of sources.