Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Synthetic Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

New tools for evaluating protein tyrosine sulphation: Tyrosyl Protein Sulphotransferases (TPSTs) are novel targets for RAF protein kinase inhibitors

Protein tyrosine sulphation is a post-translational modification (PTM) best known for regulating extracellular protein-protein interactions. Tyrosine sulphation is catalysed by two Golgi-resident enzymes termed Tyrosyl Protein Sulpho Transferases (TPSTs) 1 and 2, which transfer sulphate from the co-factor PAPS (3-phosphoadenosine 5-phosphosulphate) to a context-dependent tyrosine in a protein substrate. A lack of quantitative tyrosine sulphation assays has hampered the development of chemical biology approaches for the identification of small molecule inhibitors of tyrosine sulphation. In this paper, we describe the development of a non-radioactive mobility-based enzymatic assay for TPST1 and TPST2, through which the tyrosine sulphation of synthetic fluorescent peptides can be rapidly quantified. We exploit ligand binding and inhibitor screens to uncover a susceptibility of TPST1 and 2 to different classes of small molecules, including the anti-angiogenic compound suramin and the kinase inhibitor rottlerin. By screening the Published Kinase Inhibitor Set (PKIS), we identified oxindole-based inhibitors of the Ser/Thr kinase RAF as low micromolar inhibitors of TPST1/2. Interestingly, unrelated RAF inhibitors, exemplified by the dual BRAF/VEGFR2 inhibitor RAF265, were also TPST inhibitors in vitro. We propose that target-validated protein kinase inhibitors could be repurposed, or redesigned, as more-specific TPST inhibitors to help evaluate the sulphotyrosyl proteome. Finally, we speculate that mechanistic inhibition of cellular tyrosine sulphation might be relevant to some of the phenotypes observed in cells exposed to anionic TPST ligands and RAF protein kinase inhibitors.\n\nSUMMARY STATEMENTWe develop new assays to quantify tyrosine sulphation by the human tyrosine sulphotransferases TPST1 and 2. TPST1 and 2 catalytic activities are inhibited by protein kinase inhibitors, suggesting new starting points to synthesise (or repurpose) small molecule compounds to evaluate biological TPST using chemical biology.

biochemistry

Discrete Distributional Differential Expression (D3E) - A Tool for Gene Expression Analysis of Single-cell RNA-seq Data

The advent of high throughput RNA-seq at the single-cell level has opened up new opportunities to elucidate the heterogeneity of gene expression. One of the most widespread applications of RNA-seq is to identify genes which are differentially expressed between two experimental conditions. Here, we present a discrete, distributional method for differential gene expression (D3E), a novel algorithm specifically designed for single-cell RNA-seq data. We use synthetic data to evaluate D3E, demonstrating that it can detect changes in expression, even when the mean level remains unchanged. Since D3E is based on an analytically tractable stochastic model, it provides additional biological insights by quantifying biologically meaningful properties, such as the average burst size and frequency. We use D3E to investigate experimental data, and with the help of the underlying model, we directly test hypotheses about the driving mechanism behind changes in gene expression.

Bioinformatics

Identifying Network Perturbation in Cancer

We present a computational framework, called DISCERN (DIfferential SparsE Regulatory Network), to identify informative topological changes in gene-regulator dependence networks inferred on the basis of mRNA expression datasets within distinct biological states. DISCERN takes two expression datasets as input: an expression dataset of diseased tissues from patients with a disease of interest and another expression dataset from matching normal tissues. DISCERN estimates the extent to which each gene is perturbed - having distinct regulator connectivity in the inferred gene-regulatory dependencies between the disease and normal conditions. This approach has distinct advantages over existing methods. First, DISCERN infers conditional dependencies between candidate regulators and genes, where conditional dependence relationships discriminate the evidence for direct interactions from indirect interactions more precisely than pairwise correlation. Second, DISCERN uses a new likelihood-based scoring function to alleviate concerns about accuracy of the specific edges inferred in a particular network. DISCERN identifies perturbed genes more accurately in synthetic data than existing methods to identify perturbed genes between distinct states. In expression datasets from patients with acute myeloid leukemia (AML), breast cancer and lung cancer, genes with high DISCERN scores in each cancer are enriched for known tumor drivers, genes associated with the biological processes known to be important in the disease, and genes associated with patient prognosis, in the respective cancer. Finally, we show that DISCERN can uncover potential mechanisms underlying network perturbation by explaining observed epigenomic activity patterns in cancer and normal tissue types more accurately than alternative methods, based on the available epigenomic from the ENCODE project.

Genomics

Discovery of an autoimmunity-associated IL2RA enhancer by unbiased targeting of transcriptional activation

The majority of genetic variants associated with common human diseases map to enhancers, non-coding elements that shape cell type-specific transcriptional programs and responses to specific extracellular cues 1-3. In order to understand the mechanisms by which non-coding genetic variation contributes to disease, systematic mapping of functional enhancers and their biological contexts is required. Here, we develop an unbiased discovery platform that can identify enhancers for a target gene without prior knowledge of their native functional context. We used tiled CRISPR activation (CRISPRa) to synthetically recruit transcription factors to sites across large genomic regions (>100 kilobases) surrounding two key autoimmunity risk loci, CD69 and IL2RA (interleukin-2 receptor alpha; CD25). We identified several CRISPRa responsive elements (CaREs) with stimulation-dependent enhancer activity, including an IL2RA enhancer that harbors an autoimmunity risk variant. Using engineered mouse models and genome editing of human primary T cells, we found that sequence perturbation of the disease-associated IL2RA enhancer does not block IL2RA expression, but rather delays the timing of gene activation in response to specific extracellular signals. This work develops an approach to rapidly identify functional enhancers within non-coding regions, decodes a key human autoimmunity association, and suggests a general mechanism by which genetic variation can cause immune dysfunction.

genetics

A synthetic, three-dimensional bone marrow hydrogel

Three-dimensional (3D) synthetic hydrogels have recently emerged as desirable in vitro cell culture platforms capable of representing the extracellular geometry, elasticity, and water content of tissue in a tunable fashion. However, they are critically limited in their biological functionality. Hydrogels are typically decorated with a scant 1-3 peptide moieties to direct cell behavior, which vastly underrepresents the proteins found in the extracellular matrix (ECM) of real tissues. Further, peptides chosen are ubiquitous in ECM, and are not derived from specific proteins. We developed an approach to incorporate the protein complexity of specific tissues into the design of biomaterials, and created a hydrogel with the elasticity of marrow, and 20 marrow-specific cell-instructive peptides. Compared to generic PEG hydrogels, our marrow-inspired hydrogel improves stem cell differentiation and proliferation. We propose this tissue-centric approach as the next generation of 3D hydrogel design for applications in tissue engineering.

bioengineering

Uncovering the interplay between epigenome editing efficiency and sequence context using a novel inducible targeting system

Expanded CAG/CTG repeat disorders affect over 1 in 2500 individuals worldwide. Potential therapeutic avenues include gene silencing and modulation of repeat instability. However, there are major mechanistic gaps in our understanding of these processes, which prevent the rational design of an efficient treatment. To address this, we developed a novel system, ParB/ANCHOR-mediated Inducible Targeting (PInT), in which any protein can be recruited at will to a GFP reporter containing an expanded CAG/CTG repeat. Using PInT, we found no evidence that the histone deacetylase HDAC5 or the DNA methyltransferase DNMT1 modulate repeat instability upon targeting to the expanded repeat, suggesting that their effect is independent of local chromatin structure. Unexpectedly, we found that expanded CAG/CTG repeats reduce the effectiveness of gene silencing mediated by HDAC5 or DNMT1 targeting. The repeat-length effect in gene silencing by HDAC5 was abolished by a small molecule inhibitor of HDAC3. Our results have important implications on the design of epigenome editing approaches for expanded CAG/CTG repeat disorders. PInT is a versatile synthetic system to study the effect of any sequence of interest on epigenome editing.

molecular biology

Topological Reinforcement as a Principle of Modularity Emergence in Brain Networks

Modularity is a ubiquitous topological feature of structural brain networks at various scales. While a variety of potential mechanisms have been proposed, the fundamental principles by which modularity emerges in neural networks remain elusive. We tackle this question with a plasticity model of neural networks derived from a purely topological perspective. Our topological reinforcement model acts enhancing the topological overlap between nodes, iteratively connecting a randomly selected node to a non-neighbor with the highest topological overlap, while pruning another network link at random. This rule reliably evolves synthetic random networks toward a modular architecture. Such final modular structure reflects initial proto-modules, thus allowing to predict the modules of the evolved graph. Subsequently, we show that this topological selection principle might be biologically implemented as a Hebbian rule. Concretely, we explore a simple model of excitable dynamics, where the plasticity rule acts based on the functional connectivity between nodes represented by co-activations. Results produced by the activity-based model are consistent with the ones from the purely topological rule, showing a consistent final network configuration. Our findings suggest that the selective reinforcement of topological overlap may be a fundamental mechanism by which brain networks evolve toward modular structure.

neuroscience

Identification of novel transcriptional regulators of PKA subunits in Saccharomyces cerevisiae by Quantitative Promoter–Reporter Screening

The cAMP dependent protein kinase (PKA) signaling is a broad specificity pathway that plays important roles in the transduction of environmental signals triggering precise physiological responses. cAMP-signal transduction specificity is achieved and controlled at several levels. The Saccharomyces cereviciae PKA holoenzyme consists of two catalytic subunits encoded by TPK1, TPK2 and TPK3 genes, and two regulatory subunits encoded by BCY1 gene. In this work we studied the activity of these gene promoters using a reporter-synthetic genetic array screen, with the goal of identifying novel regulators of PKA subunits expression. Gene ontology (GO) analysis of the regulators identified showed that these regulators were enriched for annotations associated with roles in several GO biological process, as lipid and phosphate metabolism or transcription regulation and regulate all or some of the four promoters. Further characterization of the effect of these pathways on promoter activity and mRNA levels pointed to inositol, inositol polyphosphates, choline and phosphate as novel upstream signals that regulate transcription of PKA subunit genes. In addition, within each category there are genes that regulate only one of the promoters and genes that regulate more than one of them at the same time. These results support the role of transcription regulation of each PKA subunit in cAMP specificity signaling. Interestingly, many of the known targets of PKA phosphorylation are associated with the identified pathways, opening the possibility of a reciprocal regulation in which PKA would be coordinating different metabolic pathways and these processes would in turn, regulate expression of the kinase subunits.

Molecular Biology

MuClone: Somatic mutation detection and classification through probabilistic integration of clonal population structure

Accurate detection and classification of somatic single nucleotide variants (SNVs) is important in defining the clonal composition of human cancers. Existing tools are prone to miss low prevalence mutations and methods for classification of mutations into clonal groups across the whole genome are underdeveloped. Increasing interest in deciphering clonal population dynamics over multiple samples in time or anatomic space from the same patient is resulting in whole genome sequence (WGS) data from phylogenetically related samples. With the access to this data, we posited that injecting clonal structure information into the inference of mutations from multiple samples would improve mutation detection.\n\nWe developed MuClone: a novel statistical framework for simultaneous detection and classification of mutations across multiple tumour samples of a patient from whole genome or exome sequencing data. The key advance lies in incorporating prior knowledge about the cellular prevalences of clones to improve the performance of detecting mutations, particularly low prevalence mutations. We evaluated MuClone through synthetic and real data from spatially sampled ovarian cancers. Results support the hypothesis that clonal information improves sensitivity in detecting somatic mutations without compromising specificity. In addition, MuClone classifies mutations across whole genomes of multiple samples into biologically meaningful groups, providing additional phylogenetic insights and enhancing the study of WGS-derived clonal dynamics.

bioinformatics

Integrating crop growth models with whole genome prediction through approximate Bayesian computation

Genomic selection, enabled by whole genome prediction (WGP) methods, is revolutionizing plant breeding. Existing WGP methods have been shown to deliver accurate predictions in the most common settings, such as prediction of across environment performance for traits with additive gene effects. However, prediction of traits with non-additive gene effects and prediction of genotype by environment interaction (GxE), continues to be challenging. Previous attempts to increase prediction accuracy for these particularly difficult tasks employed prediction methods that are purely statistical in nature. Augmenting the statistical methods with biological knowledge has been largely overlooked thus far. Crop growth models (CGMs) attempt to represent the impact of functional relationships between plant physiology and the environment in the formation of yield and similar output traits of interest. Thus, they can explain the impact of GxE and certain types of non-additive gene effects on the expressed phenotype. Approximate Bayesian computation (ABC), a novel and powerful computational procedure, allows the incorporation of CGMs directly into the estimation of whole genome marker effects in WGP. Here we provide a proof of concept study for this novel approach and demonstrate its use with synthetic data sets. We show that this novel approach can be considerably more accurate than the benchmark WGP method GBLUP in predicting performance in environments represented in the estimation set as well as in previously unobserved environments for traits determined by non-additive gene effects. We conclude that this proof of concept demonstrates that using ABC for incorporating biological knowledge in the form of CGMs into WGP is a very promising and novel approach to improving prediction accuracy for some of the most challenging scenarios in plant breeding and applied genetics.

Genetics

Synthetic mRNA expressed Cas13a mitigates RNA virus infections

The emergence of the CRISPR-Cas system as a technology has transformed our ability to modify nucleic acids. Prokaryotes evolved one member of this family, CRISPR-Cas effector, Cas13a, as an RNA-guided ribonuclease that protects them from invading bacteriophages. Here, we demonstrate that Cas13a can be programmed to target eukaryotic viral pathogens, influenza virus A (IVA) and human respiratory syncytial virus (hRSV) in human cells. We designed synthetic mRNA coding for Cas13a, which when guided by CRISPR RNAs (crRNA) to target influenza virus or hRSV RNA, significantly mitigates these infections both prophylactically, therapeutically, and over time. These data demonstrate a possible new class of synthetic mRNA-powered anti-viral interventions.\n\nOne Sentence SummarycrRNA guided Cas13a halts RNA virus infections

molecular biology

Fitness costs of noise in biochemical reaction networks and the evolutionary limits of cellular robustness

Gene expression is inherently noisy, but little is known about whether noise affects cell function or, if so, how and by how much. Here I present a theoretical framework to quantify the fitness costs of gene expression noise and identify the evolutionary and synthetic targets of noise control. I find that gene expression noise reduces fitness by slowing the average rate of nutrient uptake and protein synthesis. This is a direct consequence of the hyperbolic (Michaelis-Menten) kinetics of most biological reactions, which I show cause \"hyperbolic filtering\", a process that diminishes both the average rate and noise propagation of stochastic reactions. Interestingly, I find that transcriptional noise directly slows growth by slowing the average translation rate. Perhaps surprisingly, this is the largest fitness cost of transcriptional noise since translation strongly filters mRNA noise, making protein noise largely independent of transcriptional noise, consistent with empirical data. Translation, not transcription, then, is the primary target of protein noise control. Paradoxically, selection for protein-noise control favors increased ribosome-mRNA binding affinity, even though this increases translational bursting. However, I find that the efficacy of selection to suppress noise decays faster than linearly with increasing cell size. This predicts a stark, cell-size-mediated taxonomic divide in selection pressures for noise control: small unicellular species, including most prokaryotes, face fairly strong selection to suppress gene expression noise, whereas larger unicells, including most eukaryotes, experience extremely weak selection. I suggest that this taxonomic discrepancy in selection efficacy contributed to the evolution of greater gene-regulatory complexity in eukaryotes.\n\nARTICLE SUMMARYGene expression is a probabilistic process, resulting in random variation in mRNA and protein abundance among cells called \"noise\". Understanding how noise affects cell function is a major problem in biology. Here I present theory demonstrating that gene expression noise slows the average rate of cell division. Furthermore, by modeling stochastic gene expression with non-linearity, I identify novel mechanisms of cellular robustness. However, I find that the cost of noise, and therefore the strength of selection favoring robustness, decays faster than linearly with increasing cell size. This may help explain the vast differences in gene-regulatory complexity between prokaryotes and eukaryotes.

Evolutionary Biology

MITRE: predicting host status from microbiota time-series data

Longitudinal studies are crucial for discovering casual relationships between the microbiome and human disease. We present Microbiome Interpretable Temporal Rule Engine (MITRE), the first machine learning method specifically designed for predicting host status from microbiome time-series data. Our method maintains interpretability by learning predictive rules over automatically inferred time-periods and phylogenetically related microbes. We validate MITREs performance on semi-synthetic data, and five real datasets measuring microbiome composition over time in infant and adult cohorts. Our results demonstrate that MITRE performs on par or outperforms \"black box\" machine learning approaches, providing a powerful new tool enabling discovery of biologically interpretable relationships between microbiome and human host.

bioinformatics

Impact of spatial organization on a novel auxotrophic interaction among soil microbes

A key prerequisite to achieve a deeper understanding of microbial communities and to engineer synthetic ones is to identify the individual metabolic interactions among key species and how these interactions are affected by different environmental factors. Deciphering the physiological basis of species-species and species-environment interactions in spatially organized environment requires reductionist approaches using ecologically and functionally relevant species. To this end, we focus here on a specific defined system to study the metabolic interactions in a spatial context among a plant-beneficial endophytic fungus Serendipita indica, and the soil-dwelling model bacterium Bacillus subtilis. Focusing on the growth dynamics of S. indica under defined conditions, we identified an auxotrophy in this organism for thiamine, which is a key co-factor for essential reactions in the central carbon metabolism. We found that S. indica growth is restored in thiamine-free media, when co-cultured with B. subtilis. The success of this auxotrophic interaction, however, was dependent on the spatial and temporal organization of the system; the beneficial impact of B. subtilis was only visible when its inoculation was separated from that of S. indica either in time or space. These findings describe a key auxotrophic interaction in the soil among organisms that are shown to be important for plant ecosystem functioning, and point to the potential importance of spatial and temporal organization for the success of auxotrophic interactions. These points can be particularly important for engineering of minimal functional synthetic communities as plant-seed treatments and for vertical farming under defined conditions.

systems biology

De novo Identification of DNA Modifications Enabled by Genome-Guided Nanopore Signal Processing

Advances in nanopore sequencing technology have enabled investigation of the full catalogue of covalent DNA modifications. We present the first algorithm for the identification of modified nucleotides without the need for prior training data along with the open source software implementation, nanoraw. Nanoraw accurately assigns contiguous raw nanopore signal to genomic positions, enabling novel data visualization, and increasing power and accuracy for the discovery of covalently modified bases in native DNA. Ground truth case studies utilizing synthetically methylated DNA show the capacity to identify three distinct methylation marks, 4mC, 5mC, and 6mA, in seven distinct sequence contexts without any changes to the algorithm. We demonstrate quantitative reproducibility simultaneously identifying 5mC and 6mA in native E. coli across biological replicates processed in different labs. Finally we propose a pipeline for the comprehensive discovery of DNA modifications in any genome without a priori knowledge of their chemical identities.

bioinformatics

A distributed algorithm to maintain and repair the trail networks of arboreal ants

We study how the arboreal turtle ant (Cephalotes goniodontus) solves a fundamental computing problem: maintaining a trail network and finding alternative paths to route around broken links in the network. Turtle ants form a routing backbone of foraging trails linking several nests and temporary food sources. This species travels only in the trees, so their foraging trails are constrained to lie on a natural graph formed by overlapping branches and vines in the tangled canopy. Links between branches, however, can be ephemeral, easily destroyed by wind, rain, or animal movements. Here we report a biologically feasible distributed algorithm, parameterized using field data, that can plausibly describe how turtle ants maintain the routing backbone and find alternative paths to circumvent broken links in the backbone. We validate the ability of this probabilistic algorithm to circumvent simulated breaks in synthetic and real-world networks, and we derive an analytic explanation for why certain features are crucial to improve the algorithms success. Our proposed algorithm uses fewer computational resources than common distributed graph search algorithms, and thus may be useful in other domains, such as for swarm computing or for coordinating molecular robots.

bioinformatics

Massively parallel digital transcriptional profiling of single cells

Characterizing the transcriptome of individual cells is fundamental to understanding complex biological systems. We describe a droplet-based system that enables 3' mRNA counting of up to tens of thousands of single cells per sample. Cell encapsulation in droplets takes place in [~]6 minutes, with [~]50% cell capture efficiency, up to 8 samples at a time. The speed and efficiency allow the processing of precious samples while minimizing stress to cells. To demonstrate the system's technical performance and its applications, we collected transcriptome data from [~][1/4] million single cells across 29 samples. First, we validate the sensitivity of the system and its ability to detect rare populations using cell lines and synthetic RNAs. Then, we profile 68k peripheral blood mononuclear cells (PBMCs) to demonstrate the system's ability to characterize large immune populations. Finally, we use sequence variation in the transcriptome data to determine host and donor chimerism at single cell resolution in bone marrow mononuclear cells (BMMCs) of transplant patients. This analysis enables characterization of the complex interplay between donor and host cells and monitoring of treatment response. This high-throughput system is robust and enables characterization of diverse biological systems with single cell mRNA analysis.

Genomics

Resolving Heterogeneous Mechanical Domains via Physics-Aware Deep Clustering of Single-Molecule Force Spectroscopy Data

Many biological processes rely on mechanical forces, with protein molecules acting as key mediators. Understanding how proteins respond to mechanical stress is essential for conditions including cardiomyopathy and muscular dystrophy. Natural proteins such as dystrophin and utrophin are composed of heterogeneous folding domains with distinct mechanical properties; deciphering domain-level behavior provides insights into disease mechanisms and informs therapeutic strategies. Single-molecule force spectroscopy (SMFS) enables probing the mechanical properties of entire proteins, yet current approaches struggle to identify heterogeneous folding domains, particularly without prior knowledge. Here, we present the first automated framework to identify heterogeneous folding domains in SMFS data, applying both existing clustering methods and a novel physics-aware deep clustering architecture, LatentUnfold. LatentUnfold learns complementary latent representations from force magnitude and the force-extension physical relationship through dual autoencoders, jointly optimized for clustering assignments. We apply our framework to experimental SMFS data collected from a synthetic two-domain protein (ddFLN4-Titin I27) as well as natural protein constructs of dystrophin and utrophin, with Monte Carlo simulated datasets serving as controlled validation. For the synthetic protein, we recover mechanical properties consistent with previously reported values for each domain. For the natural proteins, we uncover two mechanically distinct domain populations - corresponding to the N-terminal domain and spectrin-like repeats - with differences in both unfolding force and contour length increase, and reveal different unfolding order between them for the first time. This work enables domain-level biological inference, overcoming prior limitations that relied on averaging and overlooked heterogeneity, thus advancing the understanding of mechanical behavior in protein unfolding.

biophysics