Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Synthetic Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

A microfluidic biodisplay

Synthetically engineered cells are powerful and potentially useful biosensors, but it remains problematic to deploy such systems due to practical difficulties and biosafety concerns. To overcome these hurdles, we developed a microfluidic device that serves as an interface between an engineered cellular system, environment, and user. We created a biodisplay consisting of 768 individually programmable biopixels and demonstrated that it can perform multiplexed, continuous sampling. The biodisplay detected 10 {micro}g/l sodium-arsenite in tap water using a research grade fluorescent microscope, and reported arsenic contamination down to 20 {micro}g/l with an easy to interpret \"skull and crossbones\" symbol detectable with a low-cost USB microscope or by eye. The biodisplay was designed to prevent release of chemical or biological material to avoid environmental contamination. The microfluidic biodisplay thus provides a practical solution for the deployment and application of engineered cellular systems.

synthetic biology

Partitioning and Enhanced Self-Assembly of Actin in Polypeptide Coacervates

Biomolecules exist and function in cellular micro-environments that control their spatial organization, local concentration and biochemical reactivity. Due to the complexity of native cytoplasm, the development of artificial bioreactors and cellular mimics to compartmentalize, concentrate and control the local physicochemical properties is of great interest. Here, we employ self-assembling polypeptide coacervates to explore the partitioning of the ubiquitous cytoskeletal protein actin into liquid polymer-rich droplets. We find that actin spontaneously partitions into coacervate droplets and is enriched by up to {approx}30-fold. Actin polymerizes into micrometer-long filaments and, in contrast to the globular protein BSA, these filaments localize predominately to the droplet periphery. We observe up to a 50-fold enhancement in the actin filament assembly rate inside coacervate droplets, consistent with the enrichment of actin within the coacervate phase. Together these results suggest that coacervates can serve as a versatile platform in which to localize and enrich biomolecules to study their reactivity in physiological environments.\n\nSIGNIFICANCE STATEMENTLiving cells harbor many protein-rich membrane-less organelles, the biological functions of which are defined by compartment composition and properties. Significant differences between the physico-chemical properties of these crowded compartments and the dilute solutions in which biochemical reactions are traditionally studied pose a major challenge for understanding regulation of organelle composition and component activity. Here, we report the spontaneous partitioning and accelerated polymerization of the cytoskeletal protein actin inside model polypeptide coacervates as a proof-of-concept demonstration of coacervates as bioreactors for studying biomolecular reactions in cell-like environments. Our work introduces exciting avenues for the use of synthetic polymers to control the physical and biological properties of bioreactors in vitro, enabling studies of biochemical reactions in cell-like micro-environments.

biophysics

Improvement of the memory function by negative autoregulation of a mutual repression network in a stochastic environment

BackgroundCellular memory is a ubiquitous function of biological systems. By generating a sustained response to a transient inductive stimulus, often due to bistability, memory is central to the robust control of many important biological processes. However, our understanding of the origins of cellular memory remains incomplete. Stochastic fluctuations that are inherent to most biological systems have been shown to hamper memory function. Yet, how stochasticity changes the behavior of genetic circuits is generally not clear from a deterministic analysis of the network alone. Here, we apply deterministic rate equations, stochastic simulations, and theoretical analyses of Fokker-Planck equations to investigate how intrinsic noise affects the memory function in a mutual repression network.\n\nResultsWe find that the addition of negative autoregulation improves the persistence of memory in a small gene regulatory network by reducing stochastic fluctuations. Our theoretical analyses reveal that this improved memory function stems from an increased stability of the steady states of the system. Moreover, we show how the tuning of critical network parameters can further enhance memory.\n\nConclusionsOur work illuminates the power of stochastic and theoretical approaches to understanding biological circuits, and the importance of considering stochasticity to designing synthetic circuits with memory function.

systems biology

Species tree-aware simultaneous reconstruction of gene and domain evolution

Most genes are composed of multiple domains, with a common evolutionary history, that typically perform a specific function in the resulting protein. As witnessed by many studies of key gene families, it is important to understand how domains have been duplicated, lost, transferred between genes, and rearranged. Analogously to the case of evolutionary events affecting entire genes, these domain events have large consequences for phylogenetic reconstruction and, in addition, they create considerable obstacles for gene sequence alignment algorithms, a prerequisite for phylogenetic reconstruction.\n\nWe introduce the DomainDLRS model, a hierarchical, generative probabilistic model containing three levels corresponding to species, genes, and domains, respectively. From a dated species tree, a gene tree is generated according to the DL model, which is a birth-death model generalized to occur in a dated tree. Then, from the dated gene tree, a pre-specified number of dated domain trees are generated using the DL model and the molecular clock is relaxed, effectively converting edge times to edge lengths. Finally, for each domain tree and its lengths, domain sequences are generated for the leaves based on a selected model of sequence evolution.\n\nFor this model, we present a MCMC-based inference framework called DomainDLRS that takes a dated species tree together with a multiple sequence alignment for each domain family as input and outputs an estimated posterior distribution over reconciled gene and domain trees. By requiring aligned domains rather than genes, our framework evades the problem of aligning full-length genes that have been exposed to domain duplications, in particular non-tandem domain duplications. We show that DomainDLRS performs better than MrBayes on synthetic data and that it outperforms MrBayes on biological data. We analyse several zincfinger genes and show that most domain duplications have been tandem duplications, some involving two or more domains, but non-tandem duplications have also been common.

bioinformatics

Genetic dissection of courtship song variation using the Drosophila Synthetic Population Resource

Connecting genetic variation to trait variation is a grand challenge in biology. Natural populations contain a vast reservoir of fascinating and potentially useful variation, but it is unclear if the causal alleles will generally have large enough effects for us to detect. Without knowing the effect sizes or allele frequency of typical variants, it is also unclear what methods will be most successful. Here, we use a multi-parent advanced intercross population (the Drosophila Synthetic Population Resource) to map natural variation in Drosophila courtship song traits. Most additive genetic variation in this population can be explained by a modest number of highly resolved QTL. Mapped QTL are universally multiallelic, suggesting that individual genes are \"hotspots\" of natural variation due to a small target size for major mutations and/or filtering of variation by positive or negative selection. Using quantitative complementation in randomized genetic backgrounds, we provide evidence that one causal allele is harbored in the gene Fhos, making this one of the few genes associated with behavioral variation in any taxon.

Evolutionary Biology

Identification of the primary peptide inhibitor contaminant of fibrillation and toxicity in synthetic Amyloid-β42

Understanding the pathophysiology of Alzheimer disease has relied upon the use of amyloid peptides from a variety of sources, but most predominantly synthetic peptides produced using t-butyloxycarbonyl (Boc) or 9-fluorenylmethoxycarbonyl (Fmoc) chemistry. These synthetic methods can lead to minor impurities which can have profound effects on the biological activity of amyloid peptides. Here we used a combination of cytotoxicity assays, fibrillation assays and high resolution mass spectrometry (MS) to identify impurities in synthetic amyloid preparations that inhibit both cytotoxicity and aggregation. We identify the A{beta}42{Delta}39 species as the major peptide contaminant responsible for limiting both cytotoxicity and fibrillation of the amyloid peptide. In addition, we demonstrate that the presence of this minor impurity inhibits the formation of a stable A{beta}42 dimer observable by MS in very pure peptide samples. These results highlight the critical importance of purity and provenance of amyloid peptides in Alzheimers research in particular, and biological research in general.

neuroscience

An synthetic microbial loop for modelling heterotroph-phototroph metabolic interactions

Marine ecosystems are characterized by an intricate set of interactions among their representatives. One of the most important occurs through the exchange of dissolved organic matter (DOM) provided by phototrophs and used by heterotrophic bacteria as their main carbon and energy source. This metabolic interaction represents the foundation of the entire ocean food-web.\n\nHere we have assembled a synthetic ecosystem to assist the systems-level investigation of this biological association. This was achieved building an integrated, genome-scale metabolic reconstruction using two model organisms (a diatom Phaeodactylum tricornutum and an heterotrophic bacterium, Pseudoalteromonas haloplanktis) to explore and predict their metabolic interdependencies. The model was initially analysed using a constraint-based approach (Flux Balance Analysis, FBA) and then turned into a dynamic (dFBA) model to simulate a diatom-bacteria co-culture and to study the effect of changes in growth parameters on such a system. Finally, we developed a simpler dynamic ODEs system that, fed with dFBA results, was able to qualitatively describe this synthetic ecosystem and allowed performing stochastic simulations to assess the effect of noise on the overall balance of this co-culture.\n\nWe show that our model recapitulates known metabolic cross-talks of a phototroph-heterotroph system, including mutualism and competition for inorganic ions (i.e. phosphate and sulphate). Further, the dynamic simulation predicts realistic growth rate for both the diatom and the bacterium and a steady state balance between diatom and bacterial cell concentration that matches those determined in experimental co-cultures. This steady state, however, is reached following an oscillatory trend, a behaviour that is typically observed in the presence of metabolic co-dependencies. Finally, we show that, at high diatom/bacteria cell concentration ratio, stochastic fluctuations can lead to the extinction of bacteria from the co-culture, causing the explosion of diatom population. We anticipate that the developed synthetic ecosystem will serve in the future as a basis for the generation of testable hypotheses and as a scaffold for integrating and interpreting-omics data from experimental co-cultures.

systems biology

Significantly distinct branches of hierarchical trees: A framework for statistical analysis and applications to biological data

BackgroundOne of the most common goals of hierarchical clustering is finding those branches of a tree that form quantifiably distinct data subtypes. Achieving this goal in a statistically meaningful way requires (a) a measure of distinctness of a branch and (b) a test to determine the significance of the observed measure, applicable to all branches and across multiple scales of dissimilarity.\n\nResultsWe formulate a method termed Tree Branches Evaluated Statistically for Tightness (TBEST) for identifying significantly distinct tree branches in hierarchical clusters. For each branch of the tree a measure of distinctness, or tightness, is defined as a rational function of heights, both of the branch and of its parent. A statistical procedure is then developed to determine the significance of the observed values of tightness. We test TBEST as a tool for tree-based data partitioning by applying it to five benchmark datasets, one of them synthetic and the other four each from a different area of biology. For each dataset there is a well-defined partition of the data into classes. In all test cases TBEST performs on par with or better than the existing techniques.\n\nConclusionsBased on our benchmark analysis, TBEST is a tool of choice for detection of significantly distinct branches in hierarchical trees grown from biological data. An R language implementation of the method is available from the Comprehensive R Archive Network: cran.r-project.org/web/packages/TBEST/index.html.

Bioinformatics

Genome-scale genetic interactions position the Mitotic Exit Network as a major antagonist of transient Topoisomerase II deficiency.

Topoisomerase II (Top2) is the essential protein that resolves DNA catenations. When the Top2 is inactivated, mitotic catastrophe results from massive entanglement of chromosomes. Top2 is also the target of many first-line anticancer drugs, the so-called Top2 poisons. Often, tumours become resistant to these drugs by downregulating Top2. Here, we have compared two isogenic yeast strains carrying top2 thermosensitive alleles that differ in their resistance to Top2 poisons, the broadly-used poison-sensitive top2-4 and the poison-resistant top2-5. We found that top2-5 transits through anaphase faster than top2-4. In order to define the biological importance of this difference, we performed genome-scale Synthetic Gene Array (SGA) analyses during chronic sublethal Top2 downregulation and acute, yet transient, Top2 inactivation. We find that downregulation of cell cycle progression, especially the Mitotic Exit Network (MEN), protects against Top2 deficiency. In all conditions, genetic protection was stronger in top2-5, and this correlated with destabilization of anaphase bridges by execution of MEN. We suggest that mitotic exit may be a therapeutic target to hypersensitize cancer cells carrying downregulating mutations in TOP2.

genetics

Systematic Elucidation and Validation of OncoProtein-Centric Molecular Interaction Maps

The largely incomplete and tissue-independent nature of cancer pathways represents a key limitation to the ability to elucidate mechanistic determinants of cancer phenotypes and to predict adaptive response to targeted therapy. To address these challenges, we propose replacing canonical cancer pathways with a more accurate, comprehensive, and context-specific architecture - dubbed a Protein-Centric molecular interaction Map (PC-Map) - representing modulators, effectors, and cognate binding-partners of any oncoprotein of interest. To reconstruct these complex molecular architectures de novo, we introduce a novel OncoSig algorithm. Validation of a lung adenocarcinoma specific (LUAD) KRAS-centric PC-Map recapitulated known KRAS biology and, more critically, identified a novel repertoire of proteins eliciting synthetic lethality in KRASG12D LUAD organoid cultures. Showing the generalizable nature of the algorithm, we elucidated PC-Maps for ten recurrently mutated oncoproteins, including KRAS, in distinct tumor contexts. This revealed a highly context-specific nature of cancers regulatory and signaling architectures to an unprecedented degree of resolution.

systems biology

Supervised learning on synthetic data for reverse engineering gene regulatory networks from experimental time-series

The reconstruction of gene regulatory networks from time resolved gene expression measurements is a key challenge in systems biology with applications in health and disease. While the most popular network inference methods are based on unsupervised learning approaches, supervised learning methods have proven their potential for superior reconstruction performance. However, obtaining the appropriate volume of informative training data constitutes a key limitation for the success of such methods.\n\nHere, we introduce a supervised learning approach to detect gene-gene regulation based on exclusively synthetic training data, termed surrogate learning, and show its performance for synthetic and experimental time-series. We systematically investigate different simulation configurations of biologically representative time-series of transcripts and augmentation of the data with a measurement model. We compare the resulting synthetic datasets to experimental data, and evaluate classifiers trained on them for detection of gene-gene regulation from experimental time-series. For classifiers, we consider hybrid convolutional recurrent neural networks, random forests and logistic regression, and evaluate the reconstruction performance of different simulation settings, data pre-processing and classifiers.\n\nWhen training and test time-courses are generated from the same distribution, we find that the largest tested neural network architecture achieves the best performance of 0.448 {+/-} 0.047 (mean {+/-} std) in maximally achievable F1 score over all datasets outperforming random forests by 32.4 % {+/-} 14 % (mean {+/-} std). Reconstruction performance is sensitive to discrepancies between synthetic training and test data, highlighting the importance of matching training and test data domains. For an experimental gene expression dataset from E.coli, we find that training data generated with measurement model, multi-gene perturbations, but without data standardization is best suited for training classifiers for network reconstruction from the experimental test data. We further demonstrate superiority to multiple unsupervised, state-of-the-art methods for networks comprising 20 genes of the experimental data from E.coli (average AUPR best supervised = 0.22 vs best unsupervised = 0.07).\n\nWe expect the proposed surrogate learning approach to be broadly applicable. It alleviates the requirement for large, difficult to attain volumes of experimental training data and instead relies on easily accessible synthetic data. Successful application for new experimental conditions and other data types is only limited by the automatable and scalable process of designing simulations which generate suitable synthetic data.

systems biology

Haplotype-phased synthetic long reads from short-read sequencing

Next-generation DNA sequencing has revolutionized the study of biology. However, the short read lengths of the dominant instruments complicate assembly of complex genomes and haplotype phasing of mixtures of similar sequences. Here we demonstrate a method to reconstruct the sequences of individual nucleic acid molecules up to 11.6 kilobases in length from short (150-bp) reads. We show that our method can construct 99.97%-accurate synthetic reads from bacterial, plant, and animal genomic samples, full-length mRNA sequences from human cancer cell lines, and individual HIV env gene variants from a mixture. The preparation of multiple samples can be multiplexed into a single tube, further reducing effort and cost relative to competing approaches. Our approach generates sequencing libraries in three days from less than one microgram of DNA in a single-tube format without custom equipment or specialized expertise.

Genomics

MODA: MOdule Differential Analysis for weighted gene co-expression network

1Gene co-expression network differential analysis is designed to help biologists understand gene expression patterns under different conditions. We have implemented an R package called MODA (Module Differential Analysis) for gene co-expression network differential analysis. Based on transcriptomic data, MODA can be used to estimate and construct condition-specific gene co-expression networks, and identify differentially expressed subnetworks as conserved or condition specific modules which are potentially associated with relevant biological processes. The usefulness of the method is also demonstrated by synthetic data as well as Daphnia magna gene expression data under different environmental stresses.

Bioinformatics

Probabilistic adaptation in changing microbial environments

Microbes growing in animal host environments face fluctuations that have elements of both randomness and predictability. In the mammalian gut, fluctuations in nutrient levels and other physiological parameters are structured by the animal hosts behavior, diet, health and microbiota composition. Microbial cells that are able to anticipate these fluctuations by exploiting this structure would likely gain a fitness advantage, by adapting their internal state in advance. We propose that the problem of adaptive growth in these structured changing environments can be viewed as probabilistic inference. We analyze environments that are \"meta-changing\": where there are changes in the way the environment fluctuates, governed by a mechanism unobservable to cells. We develop a dynamic Bayesian model of these environments and show that a real-time inference algorithm (particle filtering) for this model can be used as a microbial growth strategy implementable in molecular circuits. The growth strategy suggested by our model outperforms heuristic strategies, and points to a class of algorithms that could support real-time probabilistic inference in natural or synthetic cellular circuits.

Systems Biology

A Theory That Predicts Behaviors Of Disordered Cytoskeletal Networks

Morphogenesis in animal tissues is largely driven by tensions of actomyosin networks, generated by an active contractile process that can be reconstituted in vitro. Although the network components and their properties are known, the requirements for contractility are still poorly understood. Here, we describe a theory that predicts whether an isotropic network will contract, expand, or conserve its dimensions. This analytical theory correctly predicts the behavior of simulated networks consisting of filaments with varying combinations of connectors, and reveals conditions under which networks of rigid filaments are either contractile or expansile. Our results suggest that pulsatility is an intrinsic behavior of contractile networks if the filaments are not stable but turn over. The theory offers a unifying framework to think about mechanisms of contractions or expansion. It provides a foundation for the study of a broad range of processes involving cytoskeletal networks, and a basis for designing synthetic networks.

cell biology

Phylogeny Recapitulates Learning: Self-Optimization of Genetic Code

Learning algorithms have been proposed as a non-selective mechanism capable of creating complex adaptive systems in life. Evolutionary learning however has not been demonstrated to be a plausible cause for the origin of a specific molecular system. Here we show that genetic codes as optimal as the Standard Genetic Code (SGC) emerge readily by following a molecular analog of the Hebbs rule (\"neurons fire together, wire together\"). Specifically, error-minimizing genetic codes are obtained by maximizing the number of physio-chemically similar amino acids assigned to evolutionarily similar codons. Formulating genetic code as a Traveling Salesman Problem (TSP) with amino acids as \"cities\" and codons as \"tour positions\" and implemented with a Hopfield neural network, the unsupervised learning algorithm efficiently finds an abundance of genetic codes that are more error-minimizing than SGC. Drawing evidence from molecular phylogenies of contemporary tRNAs and aminoacyl-tRNA synthetases, we show that co-diversification between gene sequences and gene functions, which cumulatively captures functional differences with sequence differences and creates a genomic \"memory\" of the living environment, provides the biological basis for the Hebbian learning algorithm. Like the Hebbs rule, the locally acting phylogenetic learning rule, which may simply be stated as increasing phylogenetic divergence for increasing functional difference, could lead to complex and robust life systems. Natural selection, while essential for maintaining gene function, is not necessary to act at system levels. For molecular systems that are self-organizing through phylogenetic learning, the TSP model and its Hopfield network solution offer a promising framework for simulating emerging behavior, forecasting evolutionary trajectories, and designing optimal synthetic systems.

evolutionary biology

Automatic Synthesis of Anthropomorphic Pulmonary CT Phantoms

The great density and structural complexity of pulmonary vessels and airways impose limitations on the generation of accurate reference standards, which are critical in training and in the validation of image processing methods for features such as pulmonary vessel segmentation or artery-vein (AV) separations. The design of synthetic computed tomography (CT) images of the lung could overcome these difficulties by providing a database of pseudorealistic cases in a constrained and controlled scenario where each part of the image is differentiated unequivocally. This work demonstrates a complete framework to generate computational anthropomorphic CT phantoms of the human lung automatically. Starting from biological and image-based knowledge about the topology and relationships between structures, the system is able to generate synthetic pulmonary arteries, veins, and airways using iterative growth methods that can be merged into a final simulated lung with realistic features. Visual examination and quantitative measurements of intensity distributions, dispersion of structures and relationships between pulmonary air and blood flow systems show good correspondence between real and synthetic lungs.

Bioengineering

Molecularly barcoded Zika virus libraries to probe in vivo evolutionary dynamics

Defining the complex dynamics of Zika virus (ZIKV) infection in pregnancy and during transmission between vertebrate hosts and mosquito vectors is critical for a thorough understanding of viral transmission, pathogenesis, immune evasion, and potential reservoir establishment. Within-host viral diversity in ZIKV infection is low, which makes it difficult to evaluate infection dynamics. To overcome this biological hurdle, we constructed a molecularly barcoded ZIKV. This virus stock consists of a \"synthetic swarm\" whose members are genetically identical except for a run of eight consecutive degenerate codons, which creates approximately 64,000 theoretical nucleotide combinations that all encode the same amino acids. Deep sequencing this region of the ZIKV genome enables counting of individual barcode clonotypes to quantify the number and relative proportions of viral lineages present within a host. Here we used these molecularly barcoded ZIKV variants to study the dynamics of ZIKV infection in pregnant and non-pregnant macaques as well as during mosquito infection/transmission. The barcoded virus had no discernible fitness defects in vivo, and the proportions of individual barcoded virus templates remained stable throughout the duration of acute plasma viremia. ZIKV RNA also was detected in maternal plasma from a pregnant animal infected with barcoded virus for 64 days. The complexity of the virus population declined precipitously 8 days following infection of the dam, consistent with the timing of typical resolution of ZIKV in non-pregnant macaques, and remained low for the subsequent duration of viremia. Our approach showed that synthetic swarm viruses can be used to probe the composition of ZIKV populations over time in vivo to understand vertical transmission, persistent reservoirs, bottlenecks, and evolutionary dynamics.\n\nAuthor summaryUnderstanding the complex dynamics of Zika virus (ZIKV) infection during pregnancy and during transmission to and from vertebrate host and mosquito vector is critical for a thorough understanding of viral transmission, pathogenesis, immune evasion, and reservoir establishment. We sought to develop a virus model system for use in nonhuman primates and mosquitoes that allows for the genetic discrimination of molecularly cloned viruses. This \"synthetic swarm\" of viruses incorporates a molecular barcode that allows for tracking and monitoring individual viral lineages during infection. Here we infected rhesus macaques with this virus to study the dynamics of ZIKV infection in nonhuman primates as well as during mosquito infection/transmission. We found that the proportions of individual barcoded viruses remained relatively stable during acute infection in pregnant and nonpregnant animals. However, in a pregnant animal, the complexity of the virus population declined precipitously 8 days following infection, consistent with the timing of typical resolution of ZIKV in non-pregnant macaques, and remained low for the subsequent duration of viremia.

microbiology