Search bioRxiv⌕ Search

Biology subjects

McMillen, D. R.

Publications and source records attributed to McMillen, D. R..

4 recordsLinked to original sources

A trainable language model for modulating translation rates in non-model organisms by generating upstream untranslated region sequence libraries

Tuning protein expression in non-model organisms is often constrained by the lack of validated genetic parts and predictive design tools. Translational tuning through the modulation of upstream untranslated regions (5'-UTRs) offers a potentially organism-agnostic route, but existing methods typically rely on mechanistic assumptions, prior knowledge that may not be available in non-model contexts, or the screening of sequence libraries. Here, we present a simple generative approach for creating synthetic 5'-UTR libraries based solely on the genomic sequence statistics of any desired organism. The method uses a sliding-window n-gram language model applied to native 5'-UTR sequences to produce novel sequences that preserve organism-specific base distributions and motifs without hard-coding specific motifs or mechanistic rules into inflexible statistical templates. We have applied this approach to the model bacterium Escherichia coli and the non-model probiotic Limosilactobacillus reuteri. Libraries of approximately 1,000 sequences were generated for each organism, from which about 100 unique sequences were experimentally tested for translation of a fluorescent reporter protein. In both organisms, the synthetic libraries yielded a broad range of translation levels from this relatively small number of tested variants. Sequences derived from an organisms own genomic statistics generally performed better in that organism than sequences derived from the other species. Correlations of individual sequence performance across the two species were weak, and thermodynamic predictions of ribosome binding strength showed very little predictive power, especially in the non-model L. reuteri. The results demonstrate that simple statistical language model approaches applied to genomic data can generate functional translational regulatory sequence libraries without detailed mechanistic knowledge or explicit reference to consensus motifs. The approach requires minimal computational resources, avoids reproducing native sequences, and can be readily applied to any organism with a sequenced genome. This strategy may lower technical barriers to expression tuning in non-model organisms.

synthetic biology↗

Design principles underlying nearly-homeostatic biological networks

A nearly-homeostatic system like body temperature maintenance keeps the steady state system output like internal body temperature within a narrow range, regardless of different persistent levels of environmental perturbations like external temperatures. Nearly-homeostatic systems can be implemented to guarantee performance of a therapeutic device in different patient contexts. Exploration of different near-homeostasis-supporting architectures to satisfy different design requirements is necessary, but methods for doing so in a fast and comprehensive manner remain elusive. We have identified two constrained optimization approaches to find near-homeostasis supporting architectures 10 to 100 times faster than brute-force search. Once such architectures are found, characterizing the underlying mechanisms of near-homeostasis is hindered by the very cumbersome and limited nature of traditional statistical analysis approaches. We have developed two levels of "inverse homeostasis plots", and used them to identify two novel near-homeostasis mechanisms that tolerate much higher undesired basal expression and much higher undesired first order degradation rates than previously characterized near-homeostasis mechanisms. TeaserA comprehensive toolset was developed for much faster discovery and much more comprehensive analysis of near-homeostasis-supporting architectures.

systems biology↗

A framework to enhance the Signal-to-Noise Ratio for quantitative fluorescence microscopy

Single-cell fluorescence characterization has gained much attention for studying the dynamics of individual cells in human diseases such as cancer. Despite the abundance of literature on quantitative fluorescence microscopy and its advantages in measuring cell-to-cell and spatial variation over other high-throughput instruments, it lacks a concise model that one can follow to maximize the quality of images. Here, we used the signal-to-noise ratio (SNR) model to verify camera parameters and optimize microscope settings to maximize SNR for quantitative single cell fluorescence microscopy (QSFM). We determined the microscope cameras readout noise, dark current, photon shot noise, the clock-induced charge, and validated the additive noise model for each noise source. The dark current and the clock-induced charge were both higher than reported in literature, compromising camera sensitivity. We also reduced excess background noise and improved SNR by 3-fold, by adding secondary emission and excitation filters as well as by introducing wait time in the dark before fluorescence acquisition. Additionally, our work opens new avenues for enhancing superresolution microscopy techniques such as small molecule localization microscopy (SMLM).

biophysics↗

Principles underlying implementation of nearly-homeostatic biological networks

O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=143 SRC="FIGDIR/small/646789v1_ufig1.gif" ALT="Figure 1"> View larger version (34K): org.highwire.dtl.DTLVardef@3161e2org.highwire.dtl.DTLVardef@1140a1org.highwire.dtl.DTLVardef@aa3c53org.highwire.dtl.DTLVardef@5e9ba5_HPS_FORMAT_FIGEXP M_FIG C_FIG A nearly-homeostatic biological system keeps the steady-state output (like internal body temperature) within a narrow range, regardless of different persistent levels of environmental perturbations (like external temperatures). A nearly-homeostatic system can guarantee performance of a therapeutic device in different patient contexts or can be a detector that generates a response only upon encountering an anomaly. We have developed the inverse homeostasis perspective to determine the impact of each parameter on the systems homeostatic performance, which has allowed us to vastly widen the region of implementable parameter sets supporting near-homeostasis compared to the predominant approach of minimizing the controllers integration leakiness. To implement nearly-homeostatic systems without measuring parameter values, we have used feedback-free auxiliary systems to accurately approximate steady-state response curves of key components of the feedback system and adjusted those curves to attain desired characteristics. Our approach has not only discovered a new mechanism of near-homeostasis, but also serves as a new framework to reinterpret whether the homeostatic performance observed in previously published systems arises from the mechanisms originally proposed to explain the behaviour.

synthetic biology↗