bioRxiv · 10.1101/2023.08.08.552253
Iterative Machine Learning for Classification and Discovery of Single-molecule Unfolding Trajectories from Force Spectroscopy Data
Abstract
We report the application of machine learning techniques to accelerate classification and analysis of protein unfolding trajectories from force spectroscopy data. Using kernel methods, logistic regression and triplet loss, we developed a workflow called Forced Unfolding and Supervised Iterative Online (FUSION) where a user classifies a small number of repeatable unfolding patterns encoded as image data, and a machine is tasked with identifying similar images to classify the remaining data. We tested the workflow using two case studies on a multi-domain XMod-Dockerin/Cohesin complex, validating the approach first using synthetic data generated with a Monte Carlo algorithm, and then deploying the method on experimental atomic force spectroscopy data. FUSION efficiently separated traces that passed quality filters from unusable ones, classified curves with high accuracy, and identified unfolding pathways undetected by the user. This study demonstrates the potential of machine learning to accelerate data analysis, and generate new insights in protein biophysics.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Doffini, V., Liu, H., Liu, Z., Nash, M.. 2023-08-10. Iterative Machine Learning for Classification and Discovery of Single-molecule Unfolding Trajectories from Force Spectroscopy Data. https://doi.org/10.1101/2023.08.08.552253
Cite the original work for its findings. Save a collection to share your selection of sources.