Search bioRxiv⌕ Search

Biology subjects

McLellan, J. R.

Publications and source records attributed to McLellan, J. R..

4 recordsLinked to original sources

A sequence-to-function model to predict T7 transcription rates and redesign T7 expression systems with lowered production of immunogenic RNA byproducts

T7 RNA polymerase is widely used to produce RNA using a canonical T7 promoter; however, it will also bind to low-affinity sites to generate cryptic transcription and produce RNA byproducts, which reduce full-length mRNA purity and yield. When manufacturing therapeutic RNAs for clinical applications, RNA byproducts must be removed using costly downstream purification and can cause adverse immunogenicity. To predict T7 transcription rates and reduce cryptic transcription, we designed 11588 T7 promoters and measured their mRNA levels, spanning a 6300-fold range within in vitro transcription reactions. We developed the T7 Promoter Calculator, a sequence-to-function machine learning model that predicts the T7 transcription rate on arbitrary DNA sequence across a 500-fold range with high accuracy (R2 = 0.80), accounting for both core and flanking motif sequences. We combined the model with generative design to remove low-affinity T7 sites from a therapeutic T7 expression system, resulting in a 2-fold increase in full-length mRNA purity. The automated design of T7 expression systems to remove undesired RNA byproducts increases mRNA purity and lowers downstream separation costs, while reducing adverse immunogenicity.

synthetic biology↗

A Low-Cost, High-Throughput Design-Build-Test Pipeline for Engineering Genetic Systems: Stress Testing with Complex Structural Proteins

Genetic systems engineering is constrained by high DNA synthesis costs, assembly inefficiencies, and challenges in expressing complex proteins. To address these limitations, we developed a highly parallel, low-cost pipeline for the design, assembly, and functional screening of genetic systems, which we stress-tested on highly repetitive structural proteins, including spider silk, biocements, reflectins, and talins. The integrated pipeline combines computational genetic systems design, low-cost many-plasmid DNA assembly from oligopools, automated many-to-many mapping using nanopore sequencing data, and a label-free biosensor to measure single-cell protein expression levels. We applied this pipeline to build 240 plasmids, achieving an 88% success rate (up to 2000 bp) using standard clonal isolation and 58% assembly efficiency (up to 5600 bp) without selective DNA purification, while lowering material costs by up to 24-fold. We applied the biosensor to identify genetic factors that create distinct cellular subpopulations with varying protein expression levels. Overall, the integrated pipeline will dramatically lower the cost of high-throughput synthetic biology, while demonstrating how designing genetic systems to improve build efficiency ("design for build") and directly incorporating biosensors into genetic systems ("design for test") will greatly accelerate design-build-test workflows.

synthetic biology↗

Functional Profiling of Thousands of Sequence-Diverse Protease Homologs with GROQ-seq

High-quality datasets that span broad sequence diversity are essential for understanding protein sequence-function relationships beyond local mutational landscapes. Here, we applied Growth-based Quantitative Sequencing (GROQ-seq) to measure function across an 11,722 member protease library, comprised of natural homologs and AI-shrunken variants. This library spans vast sequence diversity, with Levenshtein distances of up to 245 and a mean pairwise sequence identity of 41% to TEV protease S219V. We identified sequence-divergent TEV protease homologs that preserve function against the native TEV protease substrate. These findings reveal the robustness of protease activity across highly diverse sequences. Here, we demonstrate the aptitude of the GROQ-seq assay for screening large, diverse protein libraries for function, enabling efficient data generation at scale for training machine learning models across broad sequence landscapes.

synthetic biology↗

GROQ-seq Datasets Across Transcription Factors (LacI, RamR, VanR), T7 RNA Polymerase and TEV Protease

Predicting any proteins function from its sequence alone would be a significant breakthrough in molecular biology. Although machine learning approaches have sought to tackle this, their limited generalizability reflects the absence of sufficiently large, open, diverse, and unified datasets. To address this data gap, we developed a high-throughput experimental platform called GROQ-seq (Growth-based Quantitative Sequencing). In GROQ-seq, a proteins function can be linked to a sequencing-based readout that enables scalable characterization of large variant libraries in Escherichia coli. Here, we present pilot datasets demonstrating its performance across three distinct protein function classes: transcription factors, polymerases, and proteases. The objective of this report is to present the datasets and to provide users with a clear and transparent characterization of their properties, including both the strengths and limitations.

bioengineering↗