Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Molecular Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

The evolution of central dogma of molecular biology: a logic-based dynamic approach

It is nearly half a century past the age of the introduction of the Central Dogma (CD) of molecular biology. This biological axiom has been developed and currently appears to be all the more complex. In this study, we modified CD by adding further species to the CD information flow and mathematically expressed CD within a dynamic framework by using Boolean network based on its present-day and 1965 editions. We show that the enhancement of the Dogma not only now entails a higher level of complexity, but it also shows a higher level of robustness, thus far more consistent with the nature of biological systems. Using this mathematical modeling approach, we put forward a logic-based expression of our conceptual view of molecular biology. Finally, we show that such biological concepts can be converted into dynamic mathematical models using a logic-based approach and thus may be useful as a framework for improving static conceptual models in biology.

systems biology

25 Years of Molecular Biology Databases: A Study of Proliferation, Impact, and Maintenance

Online resources enable unfettered access to and analysis of scientific data and are considered crucial for the advancement of modern science. Despite the clear power of online data resources, including web-available databases, proliferation can be problematic due to challenges in sustainability and long-term persistence. As areas of research become increasingly dependent on access to collections of data, an understanding of the scientific communitys capacity to develop and maintain such resources is needed.\n\nThe advent of the Internet coincided with expanding adoption of database technologies in the early 1990s, and the molecular biology community was at the forefront of using online databases to broadly disseminate data. The journal Nucleic Acids Research has long published articles dedicated to the description of online databases, as either debut or update articles. Snapshots throughout the entire history of online databases can be found in the pages of Nucleic Acids Research s \"Database Issue.\" Given the prominence of the Database Issue in the molecular biology and bioinformatics communities and the relative rarity of consistent historical documentation, database articles published in Database Issues provide a particularly unique opportunity for longitudinal analysis.\n\nTo take advantage of this opportunity, the study presented here first identifies each unique database described in 3055 Nucleic Acids Research Database Issue articles published between 1991-2016 to gather a rich dataset of databases debuted during this time frame, regardless of current availability. In total, 1727 unique databases were identified and associated descriptive statistics were gathered for each, including year debuted in a Database Issue and the number of all associated Database Issue publications and accompanying citation counts. Additionally, each database identified was assessed for current availability through testing of all associated URLs published. Finally, to assess maintenance, database websites were inspected to determine the last recorded update. The resulting work allows for an examination of the overall historical trends, such as the rate of database proliferation and attrition as well as an evaluation of citation metrics and on-going database maintenance.

molecular biology

Evaluating Named-Entity Recognition approaches in plant molecular biology

Text mining research is becoming an important topic in biology with the aim to extract biological entities from scientific papers in order to extend the biological knowledge. However, few thorough studies on text mining and applications are developed for plant molecular biology data, especially rice, thus resulting a lack of datasets available to train models able to detect entities such as genes, proteins and phenotypic traits. Since there is rare benchmarks for rice, we have to face various difficulties in exploiting advanced machine learning methods for accurate analysis of rice bibliography. In this article, we developed a new training datasets (Oryzabase) as the benchmark. Then, we evaluated the performance of several current approaches to find a methodology with the best results and assigned it as the state of the art method for our own technique in the future. We applied Name Entities Recognition (NER) tagger, which is built from a Long Short Term Memory (LSTM) model, and combined with Conditional Random Fields (CRFs) to extract information of rice genes and proteins. We analyzed the performance of LSTM-CRF when applying to the Oryzabase dataset and improved the results up to 86% in F1. We found that on average, the result from LSTM-CRF is more exploitable with the new benchmark.

bioinformatics

IntelliEppi: Intelligent reaction monitoring and holistic data management system for the molecular biology lab

Daily alterations of routines and protocols create high, yet so far unmet demands for intelligent reaction monitoring, quality control and data management in molecular biology laboratories. To meet such needs, the \"internet of things\" is implemented here. We propose an approach which combines direct tracking of lab tubes, reactions and racks with a comprehensive data management system. Reagent tubes in this system are tagged with 2D data matrices or imprinted RFID-chips using a unique identification number. For each tube, individual content and all relevant information based on conducted experimental procedures are stored in an experimental data management system. This information is managed automatically but allow scientists to engage and interfere via user-friendly graphical interface. Tagged tubes are used in connection with a detectable RFID-tagged rack. We show that reaction protocols, HTS storage and complex reactions are easily planned and controlled.

bioinformatics

On a non-trivial application of Algebraic Topology to Molecular Biology

Brouwers fixed point theorem, a fundamental theorem in algebraic topology proved more than a hundred years ago, states that given any continuous map from a closed, simply connected set into itself, there is a point that is mapped unto itself. Here we point out the connection between a one-dimensional application of Brouwers fixed point theorem and a mechanism proposed to explain how extension of single-stranded DNA substrates by recombinases of the RecA superfamily facilitates significantly the search for homologous sequences on long chromosomes.

molecular biology

Universal correction of enzymatic sequence bias

Coupling molecular biology to high throughput sequencing has revolutionized the study of biology. Molecular genomics techniques are continually refined to provide higher resolution mapping of nucleic acid interactions and structure. Sequence preferences of enzymes can interfere with the accurate interpretation of these data. We developed seqOutBias to characterize enzymatic sequence bias from experimental data and scale individual sequence reads to correct intrinsic enzymatic sequence biases. SeqOutBias efficiently corrects DNase-seq, TACh-seq, ATAC-seq, MNase-seq, and PRO-seq data. We show that seqOutBias correction facilitates identification of true molecular signatures resulting from transcription factors and RNA polymerase interacting with DNA.

genomics

rxncon 2.0: a language for executable molecular systems biology

Large-scale knowledge bases and models become increasingly important to systematise and interpret empirical knowledge on cellular systems. In signalling networks, as opposed to metabolic networks, distinct modifications of and bonds between components combine into very large numbers of possible configurations, or microstates. These are essentially never measured in vivo, making explicit modelling strategies both impractical and problematic. Here, we present rxncon 2.0, the second generation rxncon language, as a tool to define signal transduction networks at the level of empirical data. By expressing both reactions and contingencies (contextual constraints on reactions) in terms of elemental states, both the combinatorial complexity and the discrepancy to empirical data can be minimised. It works as a higher-level language natural to biologists, which can be compiled into a range of graphical formats or executable models. Taken together, the rxncon language combines mechanistic precision with scalability in a composable and compilable language, that is designed for building executable knowledge bases on the molecular biology of signalling systems.

systems biology

Multiplex genome editing for synthetic biology in Vibrio natriegens

Vibrio natriegens has recently emerged as an alternative to Escherichia coli for molecular biology and biotechnology, but low-efficiency genetic tools hamper its development. Here, we uncover how to induce natural competence in V. natriegens and describe methods for multiplex genome editing by natural transformation (MuGENT). MuGENT promotes integration of large genome edits at high-efficiency on unprecedented timescales, which will extend the utility of this species for diverse applications.\n\nV. natriegens is the fastest growing organism known, with a doubling time of <10 min1,2. With broad metabolic capabilities, lack of pathogenicity, and its rapid growth rate, it is an attractive alternative to E. coli for diverse molecular biology and biotechnology applications3. Methods for classical genetic techniques have been developed for V. natriegens, but these are rel ...

synthetic biology

Current production as a rapid response expression reporter under micro-oxic and anoxic conditions

Inducible gene expression is crucial for regulating cellular processes and production of compounds within cellular pathways. Yet, inducing gene expression is only the first step to utilizing cellular processes for an applied purpose such as biosensors. Detecting when gene expression occurs is central to understanding the overall mechanism of the process as well as maximizing the process. Fluorescent proteins have been established as the primary tool for detecting gene expression in inducible systems. This study proposes electricity production as an alternate tool in reporting gene expression. Using a model organism for electricity production, Shewanella oneidensis MR-1, current was shown to be an efficient reporter for gene expression and comparable to superfolder green fluorescent protein (GFP). Through regulation of the lac operator and T7 promoter, current production was induced by isopropyl {beta}-D-1-thiogalactopyranoside (IPTG) addition. IPTG addition induced translation of GFP and the MtrB protein, which complemented a {triangleup}mtrB strain of S. oneidensis MR-1 and restored current production. This inducible system generated reproducible current in 18 minutes in both micro-oxic and anoxic conditions. These results show that current is a fast reporter for gene expression.\n\nFinancial DisclosureThe team was supported by the following departments and colleges at Michigan State University: College of Natural Science, College of Engineering, Biochemistry and Molecular Biology Department and Plant Research Laboratory. The team also received support from the DOE Great Lakes Bioenergy Research Center (DOE Office of Science BER DE-FC02-07ER64494) and startup funding from the Department of Molecular Biology and Biochemistry, Michigan State University and support from Michigan State University AgBioResearch (MICL02454) (to B.H.). This work was also supported by NSF CAREER (Award #1254238) to T.A.W. MSU Alpha Chi Sigma also supported the team.\n\nCompeting InterestsThe authors declare that no competing interests exist.\n\nEthics StatementN/A\n\nData AvailabilityAll data will be supplied upon request by the corresponding author.\n\nThis work was assessed during the iGEM/PLOS Realtime Peer Review Jamboree on 23rd February 2018 and has been revised in response to the reviewers comments.

synthetic biology

Nach is a novel ancestral subfamily of the CNC-bZIP transcription factors selected during evolution from the marine bacteria to human

All living organisms have undergone the evolutionary selection under the changing natural environments to survive as diverse life forms. All life processes including normal homeostatic development and growth into organismic bodies with distinct cellular identifications, as well as their adaptive responses to various intracellular and environmental stresses, are tightly controlled by signaling of transcriptional networks towards regulation of cognate genes by many different transcription factors. Amongst them, one of the most conserved is the basic-region leucine zipper (bZIP) family. They play vital roles essential for cell proliferation, differentiation and maintenance in complex multicellular organisms. Notably, an unresolved divergence on the evolution of bZIP proteins is addressed here. By a combination of bioinformatics with genomics and molecular biology, we have demonstrated that two of the most ancestral family members classified into BATF and Jun subgroups are originated from viruses, albeit expansion and diversification of the bZIP superfamily occur in different vertebrates. Interestingly, a specific ancestral subfamily of bZIP proteins is identified and also designated Nach (Nrf and CNC homology) on account of their highly conservativity with NF-E2 p45 subunit-related factors Nrf1/2. Further experimental evidence reveals that Nach1/2 from the marine bacteria exerts distinctive functions from Nrf1/2 in the transcriptional ability to regulate antioxidant response element (ARE)-driven cytoprotective genes. Collectively, an insight into Nach/CNC-bZIP proteins provides a better understanding of distinct biological functions between these factors selected during evolution from the marine bacteria to human.\n\nSignificanceWe identified the novel ancestral subfamily (i.e. Nach) of CNC-bZIP transcription factors with highly conservativity from marine bacteria to human. Combination of bioinformatics with genomics and molecular biology demonstrated that two of the most ancestral family members classified into BATF and Jun subgroups are originated from viruses. The Jun and CNC subfamilies also share a common origin of these bZIP proteins. Further experimental evidence reveals that Nach1/2 from the marine bacteria exerts nuance functions from human Nrf1/2 in the transcriptional ability to regulate antioxidant response element (ARE)-driven genes, responsible for the host cytoprotection against inflammation and cancer. Overall, this study is of multidisciplinary interests to provide a better understanding of distinct biological functions between Nach/CNC-bZIPs selected during evolution.

ecology

Towards a unified resource for transcriptional regulation in Escherichia coli K-12: Incorporating high throughput-generated binding data within the classic framework of regulation of initiation of transcription in RegulonDB.

Our understanding of the regulation of gene expression has been strongly benefited by the availability of high throughput technologies that enable questioning the whole genome for the binding of specific transcription factors and expression profiles. In the case of genome models, such as Escherichia coli K-12, this knowledge needs to be integrated with the legacy of accumulated genetics and molecular biology pre-genomic knowledge in order to attain deeper levels in the understanding of their biology. In spite of the several repositories and curated databases, there is no effort, nor electronic site yet, to comprehensively integrate the available knowledge from all these different sources around the regulation of gene expression of E. coli K-12. In this paper, we describe a first effort to expand RegulonDB, the database containing the rich legacy of decades of classic molecular biology experiments supporting what we know about gene regulation and operon organization in E. coli K-12, to include the genome-wide data set collections from 25 ChIP and 18 gSELEX publications, respectively, in addition to around 60 expression profiles used in their curation. Three essential features for the integration of this information coming from different methodological approaches are; first, a controlled vocabulary within an ontology for precisely defining growth conditions, second, the criteria to separate elements with enough evidence to consider them involved in gene regulation from isolated sites, and third, an expanded computational model supporting this knowledge. Altogether, this constitutes the basis for adequately gathering and enabling the comparisons and integration strongly needed to manage and access such wealth of knowledge. This version of RegulonBD is a first step toward what should become the unifying access point for current and future knowledge on gene regulation in E. coli K-12. Furthermore, this model platform and associated methodologies and criteria, can well be emulated for gathering knowledge on other microbial organisms.

systems biology

Molecular and biological characterization of an isolate of Tomato mottle mosaic virus (ToMMV) infecting tomato and other experimental hosts in a greenhouse in Valencia, Spain

Tomato is known to be a natural and experimental reservoir host for many plant viruses. In the last few years a new tobamovirus species, Tomato mottle mosaic virus (ToMMV), has been described infecting tomato and pepper plants in several countries worldwide. Upon observation of symptoms in tomato plants growing in a greenhouse in Valencia, Spain, we aimed to ascertain the etiology of the disease. Using standard molecular techniques, we first detected a positive sense single-stranded RNA virus as the probable causal agent. Next, we amplified, cloned and sequenced a ~3 kb fragment of its RNA genome which allowed us to identify the virus as a new ToMMV isolate. Through extensive assays on distinct plant species, we validated Kochs postulates and investigated the host range of the ToMMV isolate. Several plant species were locally and/or systemically infected by the virus, some of which had not been previously reported as ToMMV hosts despite they are commonly used in research greenhouses. Finally, two reliable molecular diagnostic techniques were developed and used to assess the presence of ToMMV in different plants species. We discuss the possibility that, given the high sequence homology between ToMMV and Tomato mosaic virus, the former may have been mistakenly diagnosed as the latter by serological methods.

Microbiology

Systematic assessment of GFP tag position on protein localization and growth fitness in yeast

While protein tags are ubiquitously utilized in molecular biology, they harbor the potential to interfere with functional traits of their fusion counterparts. Systematic evaluation of the effect of protein tags on localization and function would promote accurate use of tags in experimental setups. Here we examine the effect of Green Fluorescent Protein (GFP) tagging at either the N or C terminus of budding yeast proteins on localization and functionality. We use a competition-based approach to decipher the relative fitness of two strains tagged on the same protein but on opposite termini and from that infer the correct, physiological localization for each protein and the optimal position for tagging. Our study provides a first of a kind systematic assessment of the effect of tags on the functionality of proteins and provides step towards broad investigation of protein fusion libraries.\n\nHighlightsO_LIProtein tags are widely used in molecular biology although they may interfere with protein function.\nC_LIO_LIThe subcellular localization of hundreds of proteins in yeast is different when tagged at the N or the C terminus.\nC_LIO_LIA competition based assay enables systematic deciphering of correct tagging terminus for essential proteins.\nC_LIO_LIThe presented approach can be used to derive physiologically relevant tagged libraries.\nC_LI

cell biology

PrimerServer: a high-throughput primer design and specificity-checking platform

SummaryDesigning specific primers for multiple sites across the whole genome is still challenging, especially in species with complex genomes. Here we present PrimerServer, a high-throughput primer design and specificity-checking platform with both web and command-line interfaces. This platform efficiently integrates site selection, primer design, specificity checking and data presentation. In our case study, PrimerServer achieved high accuracy and a fast running speed for a large number of sites, suggesting its potential for molecular biology applications such as molecular breeding or medical testing.\n\nAvailability and ImplementationSource code for PrimerServer is available at https://github.com/billzt/PrimerServer. A demo server is freely accessible at https://primerserver.org, with all major browsers supported.\n\nContactzhangrui@caas.cn or guosandui@caas.cn

bioinformatics

Intercellular signaling network underlies biological time across multiple temporal scales

MotivationCellular, physiological and molecular processes must be organized and regulated across multiple time domains throughout the lifespan of an organism. The technological revolution in molecular biology has led to the identification of numerous genes implicated in the regulation of diverse temporal biological processes. However, it is natural to question whether there is an underlying regulatory network governing multiple timescales simultaneously.\n\nResultsUsing queries of relevant databases and literature searches, a single dense multiscale temporal regulatory network was identified involving core sets of genes that regulate circadian, cell cycle, and aging processes. The network was highly enriched for genes involved in signal transduction (P = 1.82e-82), with p53 and its regulators such as p300 and CREB binding protein forming key hubs, but also for genes involved in metabolism (P = 6.07e-127) and cellular response to stress (P = 1.56e-93). These results suggest an intertwined molecular signaling network that affects biological time across multiple temporal scales in response to environmental stimuli and available resources.\n\nContactjoshua.millstein@usc.edu\n\nSupplementary informationSupplementary data are available online.

systems biology

Coal-Miner: A Coalescent-Based Method For GWA Studies Of Quantitative Traits With Complex Evolutionary Origins

Association mapping (AM) methods are used in genome-wide association (GWA) studies to test for statistically significant associations between genotypic and phenotypic data. The genotypic and phenotypic data share common evolutionary origins - namely, the evolutionary history of sampled organisms - introducing covariance which must be distinguished from the covariance due to biological function that is of primary interest in GWA studies. A variety of methods have been introduced to perform AM while accounting for sample relatedness. However, the state of the art predominantly utilizes the simplifying assumption that sample relatedness is effectively fixed across the genome. In contrast, population genetic theory and empirical studies have shown that sample relatedness can vary greatly across different loci within a genome; this phenomena - referred to as local genealogical variation - is commonly encountered in many genomic datasets. New AM methods are needed to better account for local variation in sample relatedness within genomes.\n\nWe address this gap by introducing Coal-Miner, a new statistical AM method. The Coal-Miner algorithm takes the form of a methodological pipeline. The initial stages of Coal-Miner seek to detect candidate loci, or loci which contain putatively causal markers. Subsequent stages of Coal-Miner perform test for association using a linear mixed model with multiple effects which account for sample relatedness locally within candidate loci and globally across the entire genome.\n\nUsing synthetic and empirical datasets, we compare the statistical power and type I error control of Coal-Miner against state-of-theart AM methods. The simulation conditions reflect a variety of genomic architectures for complex traits and incorporate a range of evolutionary scenarios, each with different evolutionary processes that can generate local genealogical variation. The empirical benchmarks include a large-scale dataset that appeared in a recent high-profile publication. Across the datasets in our study, we find that Coal-Miner consistently offers comparable or typically better statistical power and type I error control compared to the state-of-art methods.\n\nCCS CONCEPTSApplied computing [->] Computational genomics; Computational biology; Molecular sequence analysis; Molecular evolution; Computational genomics; Systems biology; Bioinformatics; Population genetics;\n\nACM Reference formatHussein A. Hejase, Natalie Vande Pol, Gregory M. Bonito, Patrick P. Edger, and Kevin J. Liu. 2017. Coal-Miner: a coalescent-based method for GWA studies of quantitative traits with complex evolutionary origins. In Proceedings of ACM BCB, Boston, MA, 2017 (BCB), 10 pages. DOI: 10.475/123 4

bioinformatics

Biological screens from linear codes: theory and tools

Molecular biology increasingly relies on large screens where enormous numbers of specimens are systematically assayed in the search for a particular, rare outcome. These screens include the systematic testing of small molecules for potential drugs and testing the association between genetic variation and a phenotype of interest. While these screens are \"hypothesis-free,\" they can be wasteful; pooling the specimens and then testing the pools is more efficient. We articulate in precise mathematical ways the type of structures useful in combinatorial pooling designs so as to eliminate waste, to provide light weight, flexible, and modular designs. We show that Reed-Solomon codes, and more generally linear codes, satisfy all of these mathematical properties. We further demonstrate the power of this technique with Reed-Solomonbased biological experiments. We provide general purpose tools for experimentalists to construct and carry out practical pooling designs with rigorous guarantees for large screens.

Bioinformatics

Golden Mutagenesis: An efficient multi-site saturation mutagenesis approach by Golden Gate cloning with automated primer design

Site-directed methods for the generation of genetic diversity are essential tools in the field of directed enzyme evolution. The Golden Gate cloning technique has been proven to be an efficient tool for a variety of cloning setups. The utilization of restriction enzymes which cut outside of their recognition domain allows the assembly of multiple gene fragments obtained by PCR amplification without altering the open reading frame of the reconstituted gene. We have developed a protocol, termed Golden Muta-genesis that allows the rapid, straightforward, reliable and inexpensive construction of mutagenesis libraries. One to five amino acid positions within a coding sequence could be altered simultaneously using a protocol which can be performed within one day. To facilitate the implementation of this technique, a software library and web application for automated primer design and for the graphical evaluation of the randomization success based on the sequencing results was developed. This allows facile primer design and application of Golden Mutagenesis also for laboratories, which are not specialized in molecular biology.

molecular biology