Search bioRxiv⌕ Search

Biology subjects

Klamt, T.

Publications and source records attributed to Klamt, T..

2 recordsLinked to original sources

Predicted and experimental protein-ligand coordinates with distance labels and evaluation splits

Structure prediction tools now provide protein-ligand geometry for pairs no laboratory experiment has resolved thus far. This data is typically provided without calibrated, practically usable per-system measures of whether a given complex is correct and how much to trust it. Agreement between independently constructed methods is an established signal for this type of problem. What has been missing is a way to derive what a given level of agreement is worth. We present PLI-Parallax, which deposits predicted geometry together with experimental data needed to calibrate it. Chai-1, two Boltz-2 configurations, and the docking engine smina were run over shared inputs across an experimentally resolved crystal tier of 19,350 complexes and a corpus tier of 31,746 cross-docked pairs without experimental ground truth on predicted receptors, yielding 307,314,646 residue-to-ligand-atom distance records, which the stored coordinates allow a consumer to recompute at a cutoff of their own choosing. Their mutual agreement is fitted against observed accuracy where experimental data permits it and carried to where it does not. This way 30,567 systems carry predicted label accuracy together with a split-conformal interval. On the crystal tier that interval covers observed accuracy at the stated rate. On the corpus tier it ranks systems by expected label quality, since both distributions differ. The deposit is accompanied by 646 evaluation configurations across seven data split families, and 631 of them report how far their training and test entities separate under a two-sample test. The 906 protein accessions were partitioned to control sequence leakage, so a protein-cold split here tests generalisation across sequence space. The two tiers support evaluation against experimental data and, where this is absent, supervision weighted by how far the configurations agree.

bioinformatics↗

THEOBROMA: an aggregated open database of 1.13 million natural products with per-compound license auditing, three-tier classification, and stereochemistry-aware deduplication

AbstractNatural products remain one of the most productive sources of pharmacologically active compounds for drug discovery. Open aggregators attribute licenses at database granularity, and the consequences of that choice have become tangible as the field grows. A recent relicensing event in one constituent source (the September 2024 transition of the Natural Products Atlas to CC BY-NC 4.0) demonstrates how database-level licensing propagates across an aggregate and motivates the per-compound audit framework presented here. The same peer cohort separately leaves classification provenance and stereoisomer-family relations coarser than either layer warrants. THEOBROMA, accessible at https://theobroma.l3s.uni-hannover.de, integrates 1,132,805 natural products from 29 open sources under a per-compound license audit that resolves each compounds license tier across all attesting sources under a most-restrictive-wins rule, identifying 900,103 compounds (79.5%) under open-use licenses and providing the per-source attestation chain and resolved tier through a dedicated audit endpoint and a query- time license filter. A three-tier classification labels all but 21 compounds by the provenance of the assignment (55.3% curated source labels, 18.2% NPClassifier tool assignments, 26.5% model-inferred). Deduplication at the full 27-character InChIKey retains 1,132,805 entries across 486,032 connectivity families, exposed via a dedicated /api/stereoisomers/ endpoint and a radial-family display. Per-compound license provenance is the primary differentiator, while classification stratification and stereoisomer-family exposure add finer-grained access to two related axes, supporting license-compatible virtual screening and isomer-specific bioactivity analysis at corpus scale. As an evolving open resource, THEOBROMA pairs continuous pipeline maintenance with interactive geographic, taxonomic, and chemical-space exploration. Graphical AbstractTHEOBROMA integrates 1,132,805 natural products from 29 open sources across three axes. A per-compound license audit resolves each compound to the most restrictive tier attested by any source, giving 891,860 permissive open, 225,536 non-commercial, 8,243 public domain and 7,166 unspecified, so that 900,103 compounds (79.5%) are open-licensed. A three-tier classification records the provenance of each label as source-supplied (55.3%), tool-assigned (18.2%), or model-inferred (26.5%), leaving 21 compounds unclassified. Deduplication at the full 27-character InChIKey retains the stereochemical and protonation variants that 14-character truncation would erase, illustrated by the (E,E)- and (Z,E)-curcumin pair shown. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=80 SRC="FIGDIR/small/731585v2_ufig1.gif" ALT="Figure 1000"> View larger version (15K): org.highwire.dtl.DTLVardef@d5e1daorg.highwire.dtl.DTLVardef@1ded69eorg.highwire.dtl.DTLVardef@dc576corg.highwire.dtl.DTLVardef@1ef94b8_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗