bioRxiv · 10.64898/2026.03.15.711910
HARVEST: Unlocking the Dark Bioactivity Data of Pharmaceutical Patents via Agentic AI
Abstract
Pharmaceutical patents contain vast Structure-Activity Relationship tables documenting protein- ligand binding data. While technically public, this information remains computationally inaccessible and effectively dark, trapped in bulky documents that no existing database has systematically captured. We present HARVEST, a multi-agent large language model pipeline that autonomously extracts structured bioactivity records from USPTO patent archives at $0.11 per document. Applied to 164,877 patents, HARVEST produced 3.15 million activity records, recovering 326,342 unique scaffolds and 967 protein targets absent from BindingDB. This pipeline completed in under a week a task that would otherwise require over 55 years of continuous expert labor. Automated extraction achieves 80% agreement with human curated corpus of US patents from BindingDB, a conservative lower bound given identified errors within the reference data. We further introduce H-Bench, a structurally guaranteed held-out benchmark built from this recovered data. Evaluation of the leading open-source model Boltz-2 on H-Bench reveals a two-dimensional generalization gap: performance degrades both on novel scaffolds and on uncharacterized protein targets, exposing fundamental limitations of models trained on existing public repositories.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Shepard, V., Musin, A., Chebykina, K., Zeninskaya, N. A., Mistryukova, L., Avchaciov, K., Fedichev, P. O.. 2026-03-18. HARVEST: Unlocking the Dark Bioactivity Data of Pharmaceutical Patents via Agentic AI. https://doi.org/10.64898/2026.03.15.711910
Cite the original work for its findings. Save a collection to share your selection of sources.