Search bioRxiv⌕ Search

Biology subjects

Rieger, W. J.

Publications and source records attributed to Rieger, W. J..

3 recordsLinked to original sources

Reframing enzyme function prediction as conditional generation

Enzymes frequently exhibit promiscuous activity beyond their native roles, providing starting-points for new functions. Finding these promiscuous enzymes, especially for non-native chemical transformations, is challenging but highly valuable, as they promise novel, sustainable solutions for chemistry and biotechnology. However, current machine learning methods are poorly suited to discovering unseen chemistry as they often frame function prediction as closed set classification or a retrieval task. Here, we present Fluxion, a generative deep learning framework that learns enzymatic catalysis by modeling dynamic electron flow trajectories across the enzyme's catalytic residues. By combining both synthetic chemistry and biochemical datasets with protein language model representations, Fluxion generates multi-step electron-flow trajectories analogous to the arrow-pushing representations used to describe enzyme reaction mechanisms. Generation is conditioned on enzyme context, including the enzyme sequence, catalytic residues, substrates, and cofactors. We show that this conditioning allows Fluxion to learn enzyme-dependent regioselectivity across cytochrome P450 enzymes with different sequences shifting the predicted reaction sites for the same substrate. We then demonstrate that Fluxion's embeddings are useful for downstream tasks, such as specificity prediction on two experimental datasets, with and without finetuning. Finally, we show that Fluxion has the potential to transfer synthetic chemical logic to biology; it can generate the observed non-native product from real-world non-native directed evolution screens. Our results establish a proof of concept that generative modeling through mechanistic representations of enzymes can shift enzyme function prediction beyond static database retrieval and closed set classification to function generation. This conceptual framework provides a stepping stone towards an in silico generative method to discover non-native biocatalysts.

biochemistry↗

Squidly: Enzyme Catalytic Residue Prediction Harnessing a Biology-Informed Contrastive Learning Framework

Enzymes present a sustainable alternative to traditional chemical industries, drug synthesis, and bioremediation applications. Because catalytic residues are the key amino acids that drive enzyme function, their accurate prediction facilitates enzyme function prediction. Sequence similarity-based approaches such as BLAST are fast but require previously annotated homologs. Machine learning approaches aim to overcome this limitation; however, current gold-standard machine learning (ML)-based methods require high-quality 3D structures limiting their application to large datasets. To address these challenges, we developed Squidly, a sequence-only tool that leverages contrastive representation learning with a biology-informed, rationally designed pairing scheme to distinguish catalytic from non-catalytic residues using per-token Protein Language Model embeddings. Squidly surpasses state-of-the-art ML annotation methods in catalytic residue prediction while remaining sufficiently fast to enable wide-scale screening of databases. We ensemble Squidly with BLAST to provide an efficient tool that annotates catalytic residues with high precision and recall for both in- and out-of-distribution sequences.

bioinformatics↗

Comprehensive molecular impact mapping of common and rare variants at GWAS loci

Deep learning sequence to function models can predict the molecular effects of genetic variants, but their predictions are limited to the cell types and assays they are trained on. Here we describe DNACipher, a deep learning model that predicts the effects of genetic variants across diverse biological contexts--including those not directly measured. DNACipher takes 196 kb of genome sequences as input and imputes variant effects across 38,582 cell type-assay combinations. DNACipher generates predictions for >7 times as many contexts as Enformer, which allows for better detection of variant effects at expression quantitative trait loci (eQTLs). We also introduce DNACipher Deep Variant Impact Mapping (DVIM), a method to identify variants with molecular effects at genome-wide association study (GWAS) loci. Application of DVIM to type 1 diabetes (T1D) reduced the mean fine-mapping credible set size from 24 to 1.4 variants per signal. DVIM variants had significantly higher fine-mapping posterior probabilities, and their predicted effects were supported by single-nucleus ATAC-seq and luciferase assays. DVIM also detected 6547 rare variants with molecular effects at 96% of GWAS T1D loci, and these were enriched for associations with immune traits. In summary, DNACipher DVIM prioritises common and rare variants at GWAS loci by predicting molecular effects across a broad range of contexts.

genetics↗