bioRxiv · 10.1101/2025.11.05.686879
mdBIRCH for Fast, Scalable, Online Clustering of Molecular Dynamics Trajectories
Abstract
We present mdBIRCH, an online clustering method that adapts the BIRCH CF-tree to molecular dynamics (MD) data by applying a merge test calibrated directly to RMSD. Each arriving frame is routed to the nearest leaf microcluster and merged only if the post-merge centroid-based spread, computed from the cluster feature (CF) summaries, remains within a user-supplied threshold{tau} . This enables incremental, memory-bounded operation without constructing pairwise distance matrices, with a physically interpretable parameter controlling structural granularity. We evaluate mdBIRCH on {beta}-heptapeptide and HP35 systems and propose two practical strategies to make the threshold selection easier: (a) RMSD-anchored runs that use controlled structural edits to define interpretable operating points and (b) a blind sweep that tracks how cluster counts, occupancies, and coverage evolve with the threshold. Across both systems, increasing this threshold produces predictable consolidation into fewer, higher-occupancy states, and broader RMSD-to-centroid distributions. We further analyze the effect of the CF-tree capacity through the branching factor, quantify sensitivity to data ordering, and compare dominant-state representatives against batch clustering workflows to contextualize the resulting partitions. Finally, because decisions rely only on cluster summaries, mdBIRCH scales near-linearly with the number of frames on standard CPU hardware, offering a practical combination of speed and interpretability for large-scale trajectory analysis.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Woody Santos, J. B., Chen, L., Miranda Quintana, R. A.. 2025-11-07. mdBIRCH for Fast, Scalable, Online Clustering of Molecular Dynamics Trajectories. https://doi.org/10.1101/2025.11.05.686879
Cite the original work for its findings. Save a collection to share your selection of sources.