Search bioRxiv⌕ Search

Biology subjects

Muskal, S. M.

Publications and source records attributed to Muskal, S. M..

3 recordsLinked to original sources

Two Comparators May Be All We Need

A compound in a cell meets a spectrum of proteins drawn from many families at once, while screening most often interrogates one target at a time. Two questions asked many times in a rank ordering workflow ultimately guide decisions on what gets made and what gets counter-screened, and both are comparative: which of two targets does a compound prefer, and which of two compounds does a target prefer. We built one model for each: the target comparison over a roster of 2,279 human proteins in 34 protein families, the compound comparison over 2,079 in 32. Each model is given two chemical structures and a sequence, or two sequences and a chemical structure, and returns which member of the pair is preferred together with how firmly it holds that view. No conformational analysis, protein structure, binding site or docked pose is used. Asked which of two targets a compound prefers, the model is correct 0.75 of the time over 8,689 held-out comparisons setting two families against each other, and 0.78 of the time over 32,738 comparisons between two targets of one family, rising to 0.94 and 0.95 on the most confident third of each. Asked which of two compounds a single target prefers, it is correct 0.71 of the time over 65,725 held-out comparisons, rising to 0.96 on the most confidently held. Neither compound in any of those comparisons appeared anywhere in training. Accuracy follows the gap between the two measurements. Where they sit within half a log unit the models are right 0.57 to 0.64 of the time, and where they differ by more than two logs, 0.90 to 0.92. The compound result holds across 31 protein families, not only the best-measured one. Within the chemistry and the targets they were built on, these models rank compounds and rank targets well. Both models can be explored and downloaded at familyfoundationmodel.com. Keywords: target preference; compound preference; pairwise comparison; polypharmacology; off-target triage; ESM2; random forest; structure-free prediction; ChEMBL

bioinformatics↗

ReverseScreen.ai: Pharmacophore-Guided Reverse Screening Across the Growing Co-Complex Proteome

Biopharmaceutical companies routinely use forward screening to identify potential ligands for targets of biological consequence. Increasingly, they are also using reverse screening to proactively identify downstream off-target liabilities and new repurposing opportunities. Exhaustive reverse docking of one or a multitude of molecules across every characterized site is one approach, but the computational cost is often prohibitive. One molecule against 28,579 receptor sites takes 40.5 hours on twenty CPU cores. We present a method that puts a retrieval step in front of the docking. Every co-crystallized ligand in the Protein Data Bank is indexed by the 3-dimensional pharmacophore fingerprint it presents, together with the UniProt accessions it was solved against, and a query retrieves the sites belonging to its nearest neighbors. Fingerprints are predicted from two-dimensional structures with PharmCast, so a query is fingerprinted in 4 ms and searched against 27,797 indexed ligands in 40 ms. For a query molecule we take its 5 most similar indexed ligands and pool every protein those 5 were crystallized with, which averages 21.1 distinct proteins. Across 3,000 held-out molecules that pool contained the molecule's own known target 48.8 percent of the time. Docking those 21.1 proteins takes 108 s, against 40.5 hours for the full panel. Pooling the twenty-five most similar ligands instead gives 104.5 proteins and finds the target 60.8 percent of the time. To be indexed, a structure only has to establish which ligand sat in which protein. Holding the index to pocket-grade coordinates had been excluding whole receptor classes. Admitting X-ray at 2.5 [A] and cryo-EM at 4.0 [A] takes coverage from 3,670 target sites to 28,579. Two 2025 clinical molecules were cross-validated. Orforglipron returned the glucagon-like peptide 1 (GLP-1) receptor first of 27,797 ligands, through a non-identical analog at 0.901; daraxonrasib returned the KRAS and cyclophilin A tri-complex fourth, at 0.836. When run in batches the pipeline reverse screens about 96,000 molecules per hour on 1 core, roughly 2.3 million per day, so the retrieval step is practical even with a massive virtual library. The method is available at reversescreen.ai, which allows users the opportunity to screen molecules against the current index, returns the retrieved sites, and docks them individually. The index itself can be downloaded for confidential screening in a local environment. Keywords: reverse screening, target identification, pharmacophore fingerprint, molecular docking, polypharmacology, off-target prediction, virtual screening, repurposing

bioinformatics↗

PharmCast: rapid generation of three-dimensional pharmacophore fingerprints from two-dimensional structure without conformer generation

A three-dimensional pharmacophore fingerprint records the binding features a molecule can present. It is a description of a hand in search of a glove. Because it is defined by presented features instead of two-dimensional structure, it can identify pharmacophoric similarity between structurally distinct compounds, which is what scaffold hopping and non-obvious me-too design require. The descriptor has remained a niche tool because its cost is dominated by conformer generation. In the reference pipeline, generating 100 conformers requires 2.82 s of the 2.86 s needed to fingerprint one screening collection compound; the bit calculation requires 0.039 s. We therefore removed the conformational stage. PharmCast is a feedforward neural network that predicts all 10,549 bits of a PharmPrint ensemble fingerprint directly from a SMILES string. On the same machine, PharmCast generated pharmacophore fingerprints for two molecules and compared them in 0.584 ms, whereas the conventional conformer-based pipeline took 5.71 s. PharmCast version 10 was trained on 5,887,229 molecules drawn from a screening collection, activity-backed ChEMBL compounds from 142 to 1000 Da, and peptide loops excised from crystal structures. We evaluated 155,648 purchasable catalog compounds excluded from every training set, 139,700 activity-backed ChEMBL compounds not present in the version 10 training set, and 13,500 peptide loops reserved for testing. Median fingerprint error, Pearson r, and pairwise ranking accuracy were 0.008, 0.980, and 0.936 for screening collection chemistry; 0.016, 0.984, and 0.952 for loop peptides; and 0.027, 0.936, and 0.889 for activity-backed ChEMBL compounds. The reference calculation reproduces itself at an error of 0.006 and r of 0.995. Ensemble pharmacophore fingerprints can therefore be predicted from two-dimensional structure alone at a cost suitable for large-scale collection screening and virtual library exploration. Keywords: pharmacophore, fingerprint, scaffold hopping, surrogate model, virtual screening, applicability domain

bioinformatics↗