Search bioRxiv⌕ Search

Biology subjects

Ramdhan, P. A.

Publications and source records attributed to Ramdhan, P. A..

2 recordsLinked to original sources

GatorAffinity: Boosting Protein-Ligand Binding Affinity Prediction with Large-Scale Synthetic Structural Data

Protein-ligand binding affinity prediction is a fundamental task in computational drug discovery. Although substantial efforts have been made to enhance prediction accuracy using data-driven approaches, progress remains limited by persistent data scarcity. The widely used PDBbind dataset, for example, contains fewer than 20, 000 experimental structures with annotated binding affinities, while a vast number of affinity measurements remain underutilized due to missing structural data. Here, we investigate this untapped potential by curating more than 450, 000 synthetic protein-ligand complexes annotated with Kd and Ki values using the Boltz-1 structure prediction model. Building on this unprecedented scale of synthetic data, further augmented with over 1 million synthetic complexes from the recently released SAIR database annotated with IC50 values, we develop GatorAffinity, a geometric deep learning-based scoring function pretrained on large-scale synthetic data and fine-tuned using high-quality experimental structures from PDBbind. Extensive evaluation on a leak-proof benchmark demonstrates that GatorAffinity significantly outperforms state-of-the-art affinity prediction methods, offering superior accuracy and generalizability. Our findings show that augmenting available experimental data with synthetic complexes can effectively address the data scarcity challenge while maintaining strong predictive reliability. By releasing the pretrained GatorAffinity model and the large-scale synthetic dataset GatorAffinity-DB, we provide a scalable and reproducible foundation for affinity prediction, virtual screening, and broader structure-based drug design applications (https://github.com/AIDD-LiLab/GatorAffinity).

bioinformatics↗

MGMG: Cell Morphology-Guided Molecule Generation for Drug Discovery

Designing novel molecules with desired bioactivity remains a fundamental challenge in drug discovery. Most molecular design methods follow target-based drug discovery paradigms that rely on well-defined drug targets, thereby limiting their applicability to diseases lacking known targets or reference compounds. Here we introduce Morphology-Guided Molecule Generation (MGMG), a phenotypic drug discovery-oriented approach that integrates cellular morphological profiles from compound treatments with molecular textual descriptions without requiring molecular target information. Cell morphology offers the guidance on desired bioactivity-relevant phenotypic effects, while textual descriptions provide direct and interpretable cues about molecular structure. Leveraging complementary structural and bioactivity context, MGMG significantly enhances molecule generation performance, especially in scenarios where textual descriptions are under-informative or morphological signals are weak. MGMG can also be applied to genetic perturbations, enabling activator design from gene overexpression morphology without requiring knowledge of reference compound structure. In addition, in silico docking demonstrates that MGMG-generated molecules, despite lacking target information, exhibit binding affinities comparable to reference compounds, preserving key interactions while introducing structural diversity. Overall, MGMG jointly utilizes morphological and textual description inputs to guide molecule generation, enabling diverse, bioactivity-aware compound design in a target-agnostic fashion.

bioinformatics↗