Search bioRxiv⌕ Search

Biology subjects

Dolorfino, M. D.

Publications and source records attributed to Dolorfino, M. D..

2 recordsLinked to original sources

Assessing the Generalizability of Machine Learning and Physics Methods for DNA-Encoded Libraries

Predicting protein-ligand binding is a central challenge in computational drug discovery, and while machine learning (ML) and co-folding methods have advanced rapidly, their ability to generalize beyond training or parameterization regimes remains insufficiently understood. DNA-encoded libraries (DELs) enable ultra-large screening of billions of molecules simultaneously, providing a useful testbed for evaluating these approaches at scale. A recent NeurIPS competition revealed that even top performing ML models trained on DEL data failed at generalizing to out-of-distribution (OOD) chemical space. We investigated whether integrating structural modeling could bridge this generalization gap. We systematically assessed state-of-the-art ML, docking, and co-folding methods including Schrodinger Glide, Rosetta GALigandDock, and Boltz-2 with three biologically diverse protein targets screened against libraries containing multiple DEL synthesis formats. While ML excels in-distribution, OOD hit discrimination is dependent on both the target and ligand context, with no single method consistently dominating. These findings demonstrate that benchmark performance alone is insufficient to predict OOD performance, highlighting the need for system-dependent evaluation of binding prediction methods. We provide an open-source package for assessing protein-ligand prediction methods and analyzing high-throughput screening data: DEL-iver. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=118 SRC="FIGDIR/small/719394v2_ufig1.gif" ALT="Figure 1"> View larger version (25K): org.highwire.dtl.DTLVardef@cd8a17org.highwire.dtl.DTLVardef@252d0aorg.highwire.dtl.DTLVardef@b020f0org.highwire.dtl.DTLVardef@1429e27_HPS_FORMAT_FIGEXP M_FIG C_FIG

biophysics↗

ProteinMPNN Recovers Complex Sequence Properties of Transmembrane β-Barrels

Recent deep-learning (DL) protein design methods have been successfully applied to a range of protein design problems, including the de novo design of novel folds, protein binders, and enzymes. However, DL methods have yet to meet the challenge of de novo membrane protein (MP) and the design of complex {beta}-sheet folds. We performed a comprehensive benchmark of one DL protein sequence design method, ProteinMPNN, using transmembrane and water-soluble {beta}-barrel folds as a model, and compared the performance of ProteinMPNN to the new membrane-specific Rosetta Franklin2023 energy function. We tested the effect of input backbone refinement on ProteinMPNN performance and found that given refined and well-defined inputs, ProteinMPNN more accurately captures global sequence properties despite complex folding biophysics. It generates more diverse TMB sequences than Franklin2023 in pore-facing positions. In addition, ProteinMPNN generated TMB sequences that passed state-of-the-art in silico filters for experimental validation, suggesting that the model could be used in de novo design tasks of diverse nanopores for single-molecule sensing and sequencing. Lastly, our results indicate that the low success rate of ProteinMPNN for the design of {beta}-sheet proteins stems from backbone input accuracy rather than software limitations.

bioinformatics↗