Search bioRxiv⌕ Search

Biology subjects

McEllin, B.

Publications and source records attributed to McEllin, B..

2 recordsLinked to original sources

Structure-free, site-resolved contrastive learningextends small-molecule discovery beyond the reachof structure-based modeling

Virtual screening asks which molecules, among an enormous space of drug-like chemistry, are worth synthesizing and testing against a protein target. Most modern methods answer by building and scoring an explicit three-dimensional pose through molecular docking, or the co-folding models that now approach experimental accuracy. Building these poses presumes a well-defined pocket, and the non-orthosteric, cryptic, and intrinsically disordered sites where unexplored ligandability lies offer none. There, these methods fail to generalize. Here we present Ptarmigan-1, a contrastive model that co-embeds the residues of a protein with candidate small molecules in a shared latent space, from sequence and two-dimensional chemistry alone, and without ever constructing a pose. Engagement reduces to the proximity of precomputed embeddings. Freed from the pose, Ptarmigan-1 trains directly on chemoproteomic and bioactivity data of mixed resolution, scores a compound in ten milliseconds rather than the tens of seconds a co-folding model demands, and resolves each prediction to the residues a compound engages. On well-folded, orthosteric targets it performs comparably to a collection of co-folding and docking models, and on covalent, cryptic, and disordered sites it matches or exceeds them. It localizes reversible and covalent inhibitors to the pockets they engage, even for targets withheld from training, and screens the entire human proteome against a library of 3.4 billion compounds in under a day. By decoupling molecular recognition from structure, Ptarmigan-1 recasts virtual screening as a reusable index that continuously improves as data accumulate.

molecular biology↗

Native, Spatiotemporal Profiling of the Global Human Regulome

The regulome, comprising transcription factors, cofactors, chromatin remodelers, and other regulatory proteins, forms the core machinery by which cells interpret signals and execute gene expression programs. Despite its central role in development, disease, and drug response, the regulome remains largely uncharted at scale due to its dynamic, low-abundance, and chromatin-associated nature. Here, we present a method for scalable, regulome profiling for global, compartment-resolved quantification of native regulome proteins. By enriching DNA- and chromatin-associated proteins and profiling them using high-throughput, label-free DIA mass spectrometry, regulome profiling captures chromatin-associated proteins across 36 human cell lines and thousands of perturbations. The resulting Regulome Atlas recovers nearly 60% of known human transcription factors, reveals lineage-specific TF localization, and distinguishes active nuclear engagement from latent, unbound states. We demonstrate that regulome profiles resolve acute immune pathway activation prior to transcriptional changes, identify previously unrecognized drug-induced regulome responses, and enable proteome-scale readouts of compound target engagement and complex remodeling. This work establishes a foundational resource for decoding the regulatory proteome and provides a blueprint for integrating regulome data into next-generation models of cellular behavior.

cell biology↗