Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Molecular Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,153 records · Page 64Linked to original sources

Deciphering the genic basis of environmental adaptations by simultaneous forward and reverse genetics in Saccharomyces cerevisiae

The budding yeast Saccharomyces cerevisiae is the best studied eukaryote in molecular and cell biology, but its utility for understanding the genetic basis of natural phenotypic variation is limited by the inefficiency of association mapping owing to strong and complex population structure. To facilitate association mapping, we analyzed 190 high-quality genomes of diverse strains, including 85 newly sequenced ones, to uncover yeasts population structure that varies substantially among genomic regions. We identified 181 yeast genes that are absent from the reference genome and demonstrated their expression and role in important functions such as drug resistance. We then simultaneously measured the growth rates of over 4500 lab strains each deficient of a nonessential gene and 81 natural strains across multiple environments using unique DNA barcode present in each strain. We combined the genome-wide reverse genetic information with genome-wide association analysis to determine potential genomic regions of importance to environmental adaptations, and for a subset experimentally validated their role by reciprocal hemizygosity tests. The resources provided permit efficient and reliable association mapping in yeast and significantly enhances its value as a model for understanding the genetic mechanisms of phenotypic polymorphism and evolution.

genomics

A systematic comparison of error correction enzymes by next-generation sequencing

Gene synthesis, the process of assembling gene-length fragments from shorter groups of oligonucleotides (oligos), is becoming an increasingly important tool in molecular and synthetic biology. The length, quality, and cost of gene synthesis is limited by errors produced during oligo synthesis and subsequent assembly. Enzymatic error correction methods are cost-effective means to ameliorate errors in gene synthesis. Previous analyses of these methods relied on cloning and Sanger sequencing to evaluate their efficiencies, limiting quantitative assessment and throughput. Here we develop a method to quantify errors in synthetic DNA by next-generation sequencing. We analyzed errors in a model gene assembly and systematically compared six different error correction enzymes across 11 conditions. We find that ErrASE and T7 Endonuclease I are the most effective at decreasing average error rates (up to 5.8-fold relative to the input), whereas MutS is the best for increasing the number of perfect assemblies (up to 25.2-fold). We are able to quantify differential specificities such as ErrASE preferentially corrects C/G [->] G/C transversions whereas T7 Endonuclease I preferentially corrects A/T [->] T/A transversions. More generally, this experimental and computational pipeline is a fast, scalable, and extensible way to analyze errors in gene assemblies, to profile error correction methods, and to benchmark DNA synthesis methods.

synthetic biology

Uncovering Robust Patterns of MicroRNA Co-Expression across Cancers using Bayesian Relevance Networks

Co-expression networks have long been used as a tool for investigating the molecular circuitry governing biological systems. However, most algorithms for constructing co-expression networks were developed in the microarray era, before high-throughput sequencing--with its unique statistical properties--became the norm for expression measurement. Here we develop Bayesian Relevance Networks, an algorithm that uses Bayesian reasoning about expression levels to account for the differing levels of uncertainty in expression measurements between highly- and lowly-expressed entities, and between samples with different sequencing depths. It combines data from groups of samples (e.g., replicates) to estimate group expression levels and confidence ranges. It then computes uncertainty-moderated estimates of cross-group correlations between entities, and uses permutation testing to assess their statistical significance. Using large scale miRNA data from The Cancer Genome Atlas, we show that our Bayesian update of the classical Relevance Networks algorithm provides improved reproducibility in co-expression estimates and lower false discovery rates in the resulting co-expression networks. Software is available at www.perkinslab.ca/Software.html.

bioinformatics

Predicting DNA Hybridization Kinetics from Sequence

Hybridization is a key molecular process in biology and biotechnology, but to date there is no predictive model for accurately determining hybridization rate constants based on sequence information. To approach this problem systematically, we first performed 210 fluorescence kinetics experiments to observe the hybridization kinetics of 100 different DNA target and probe pairs (subsequences of the CYCS and VEGF genes) at temperatures ranging from 28 {degrees}C to 55 {degrees}C. Next, we rationally designed 38 features computable based on sequence, each feature individually correlated with hybridization kinetics. These features are used in our implementation of a weighted neighbor voting (WNV) algorithm, in which the hybridization rate constant of an unknown sequence is predicted based on similarity reactions with known rate constants (a.k.a. labeled instances). Automated feature selection and weighting optimization resulted in a final 6-feature WNV model, which can predict hybridization rate constants of new sequences to within a factor of 2 with {approx}74% accuracy and within a factor of 3 with {approx}92% accuracy, based on leave-one-out cross-validation. Predictive understanding of hybridization kinetics allows more efficient design of nucleic acid probes, for example in allowing sparse hybrid-capture panels to more quickly and economically enrich desired regions from genomic DNA.

biophysics

Cell cycle time series gene expression data encoded as cyclic attractors in Hopfield systems

Modern time series gene expression and other omics data sets have enabled unprecedented resolution of the dynamics of cellular processes such as cell cycle and response to pharmaceutical compounds. In anticipation of the proliferation of time series data sets in the near future, we use the Hopfield model, a recurrent neural network based on spin glasses, to model the dynamics of cell cycle in HeLa (human cervical cancer) and S. cerevisiae cells. We study some of the rich dynamical properties of these cyclic Hopfield systems, including the ability of populations of simulated cells to recreate experimental expression data and the effects of noise on the dynamics. Next, we use a genetic algorithm to identify sets of genes which, when selectively inhibited by local external fields representing gene silencing compounds such as kinase inhibitors, disrupt the encoded cell cycle. We find, for example, that inhibiting the set of four kinases BRD4, MAPK1, NEK7, and YES1 in HeLa cells causes simulated cells to accumulate in the M phase. Finally, we suggest possible improvements and extensions to our model.\n\nAuthor SummaryCell cycle - the process in which a parent cell replicates its DNA and divides into two daughter cells - is an upregulated process in many forms of cancer. Identifying gene inhibition targets to regulate cell cycle is important to the development of effective therapies. Although modern high throughput techniques offer unprecedented resolution of the molecular details of biological processes like cell cycle, analyzing the vast quantities of the resulting experimental data and extracting actionable information remains a formidable task. Here, we create a dynamical model of the process of cell cycle using the Hopfield model (a type of recurrent neural network) and gene expression data from human cervical cancer cells and yeast cells. We find that the model recreates the oscillations observed in experimental data. Tuning the level of noise (representing the inherent randomness in gene expression and regulation) to the \"edge of chaos\" is crucial for the proper behavior of the system. We then use this model to identify potential gene targets for disrupting the process of cell cycle. This method could be applied to other time series data sets and used to predict the effects of untested targeted perturbations.

systems biology

Ultra-fast super-resolution imaging of biomolecular mobility in tissues

Super-resolution techniques have addressed many biological questions, yet molecular quantification at rapid timescales in live tissues remains challenging. We developed a light microscopy system capable of sub-millisecond sampling to characterize molecular diffusion in heterogeneous aqueous environments comparable to interstitial regions between cells in tissues. We demonstrate our technique with super-resolution tracking of fluorescently labelled chemokine molecules in a collagen matrix and ex vivo lymph node tissue sections, outperforming competing methods.

biophysics

Weakly-bound Dimers that Underlie the Crystal Nucleation Precursors in Lysozyme Solutions

Protein crystallization is central to understanding of molecular structure in biology, a vital part of processes in the pharmaceutical industry, and a crucial component of numerous disease pathologies. Crystallization starts with nucleation and how nucleation proceeds determines the crystallization rate and essential properties of the resulting crystal population. Recent results with several proteins indicate that crystals nucleate within preformed mesoscopic protein-rich clusters. The origin of the mesoscopic clusters is poorly understood. In the case of lysozyme, a common model of protein biophysics, earlier findings suggest that clusters exist owing to the dynamics of formation and decay of weakly-bound transient dimers. Here we present evidence of a weakly bound lysozyme dimer in solutions of this protein. We employ two electrospray mass spectrometry techniques, a combined ion mobility separation mass spectrometry and a high-resolution implementation. To enhance the weak but statistically-significant dimer signal we develop a method based on the residuals between the maxima of the isotope peaks in Fourier space and their Gaussian envelope. We demonstrate that these procedures sensitively detect the presence of a non-covalently bound dimer and distinguish its signal from other polypeptides, noise, and sampling artefacts. These findings contribute essential elements of the crystal nucleation mechanism of lysozyme and other proteins and suggest pathways to control nucleation and crystallization by enhancing or suppressing weak oligomerization.

biophysics

Monitoring changes in the Gene Ontology and their impact on genomic data analysis

The Gene Ontology (GO) is one of the most widely used resources in molecular and cellular biology, largely through the use of \"enrichment analysis\". To facilitate informed use of GO, we present GOTrack (https://gotrack.msl.ubc.ca), which provides access to historical records and trends in the Gene Ontology and GO annotations (GOA). GOTrack gives users access to gene- and term-level information on annotations for nine model organisms as well as an interactive tool that measures the stability of enrichment results over time for user-provided \"hit lists\" of genes. To document the effects of GO evolution on enrichment, we analyzed over 2500 published hit lists of human genes (most over 9 years old). 53% of hit lists were considered to yield significantly stable enrichment results. Because stability is far from assured for any individual hit list, GOTrack can lead to more informed and cautious application of GO to genomics research.

bioinformatics

Humoral immune response to adenovirus induce tolerogenic bystander dendritic cells that promote generation of regulatory T cells

Following repeated encounters with adenoviruses most of us develop robust humoral and cellular immune responses that are thought to act together to combat ongoing and subsequent infections. Yet in spite of robust immune responses, adenoviruses establish subclinical persistent infections that can last for decades. While adenovirus persistence pose minimal risk in B-cell compromised individuals, if T-cell immunity is severely compromised, reactivation of latent adenoviruses can be life threatening. This dichotomy led us to ask how anti-adenovirus antibodies influence adenovirus-specific T-cell immunity. Using primary human blood cells, transcriptome and secretome profiling, and pharmacological, biochemical, genetic, molecular, and cell biological approaches, we initially found that healthy adults harbor adenovirus-specific regulatory T cells (Tregs). As peripherally induced Tregs are generated by tolerogenic dendritic cells (DCs), we then addressed how tolerogenic DCs could be created. Here, we demonstrate that DCs that take up immunoglobulin-complexed (IC)-adenoviruses create an environment that causes bystander DCs to become tolerogenic. These adenovirus antigen-loaded tolerogenic DCs can drive naive T cells to mature into adenovirus-specific Tregs. Our results may provide ways to improve antiviral therapy and/or pre-screening high-risk individuals undergoing immunosuppression.\n\nAuthor summaryWhile numerous studies have addressed the cellular and humoral response to primary virus encounters, relatively little is known about the interplay between persistent infections, neutralizing antibodies, antigen-presenting cells, and the T-cell response. Our studies suggests that if adenovirus-antibody complexes are taken up by professional antigen-presenting cells (dendritic cells), the DCs generate an environment that causes bystander dendritic cells to become tolerogenic. These tolerogenic dendritic cells favors the creation of adenovirus-specific regulatory T cells. While this pathway likely favors pathogen survival, there may be advantages for the host also.

microbiology

BALLI: Bartlett-Adjusted Likelihood-based LInear Model Approach for Identifying Differentially Expressed Gene with RNA-seq Data

MotivationTranscriptomic profiles can improve our understanding of the phenotypic molecular basis of biological research, and many statistical methods have been proposed to identify differentially expressed genes under two or more conditions with RNA-seq data. However, statistical analyses with RNA-seq data often suffer from small sample sizes, and global variance estimates of RNA expression levels have been utilized as prior distributions for gene-specific variance estimates, making it difficult to generalize the methods to more complicated settings. We herein proposed a Bartlett-Adjusted Likelihood based LInear mixed model approach (BALLI) to analyze more complicated RNA-seq data. The proposed method estimates the technical and biological variances with a linear mixed effect model, with and without adjusting small sample bias using Bartletts corrections.\n\nResultsWe conducted extensive simulations to compare the performance of BALLI with those of existing approaches (edgeR, DESeq2, and voom). Results from the simulation studies showed that BALLI correctly controlled the type-1 error rates at the various nominal significance levels, and produced better statistical power and precision estimates than those of other competing methods in various scenarios. Furthermore, BALLI was robust to variation of library size. It was also successfully applied to Holstein milk yield data, illustrating its practical value.\n\nAvailability and ImplementationBALLI is implemented as R package and freely available at http://healthstat.snu.ac.kr/software/balli/.\n\nContactwon1@snu.ac.kr\n\nSupplementary InformationSupplementary data are available at Bioinformatics online

bioinformatics

A Next Generation Clustering Tool Enables Identification of Functional Cancer Subtypes with Associated Biological Phenotypes

One of the major challenges faced in defining clinically applicable and homogeneous molecular tumor subtypes is assigning biological and/or clinical interpretations to etiological (intrinsic) subtypes. The conventional approach involves at least three steps: Firstly, identify subtypes using unsupervised clustering of patient tumours with molecular (etiological) profiles; secondly associate the subtypes with clinical or phenotypic information (covariates) to infer some biological meaning to the redefined subtypes; and thirdly, identify clinically relevant biomarkers associated with the subtypes. Here, we report the implementation of a tool, phenotype mapping (phenMap), which combines these three steps to define functional subtypes with associated phenotypic information and molecular signatures. phenMap models meta (unobserved) variables as a function of covariates to expose any underlying clustering structure within the data and discover associations between subtypes and phenotypes. We demonstrate how this tool can more avidly identify functional subtypes that are an improvement over already existing etiological subtypes by analysing published breast cancer gene expression data.

genomics

DeDaL: Cytoscape 3.0 app for producing and morphing data-driven and structure-driven network layouts

BackgroundVisualization and analysis of molecular profiling data together with biological networks are able to provide new mechanistical insights into biological functions. Currently, high-throughput data are usually visualized on top of predefined network layouts which are not always adapted to a given data analysis task. We developed a Cytoscape app which allows to construct biological network layouts based on the data from molecular profiles imported as values of nodes attributes.\n\nResultsDeDaL is a Cytoscape 3.0 app which uses linear and non-linear algorithms of dimension reduction to produce data-driven network layouts based on multidimensional data (typically gene expression). DeDaL implements several data pre-processing and layout post-processing steps such as continuous morphing between two arbitrary network layouts and aligning one network layout with respect to another one by rotating and mirroring. Combining these possibilities facilitates creating insightful network layouts representing both structural network features and the correlation patterns in multivariate data.\n\nConclusionsDeDaL is the first method allowing to construct biological network layouts from high-throughput data. DeDaL is freely available for downloading together with step-by-step tutorial at http://bioinfo-out.curie.fr/projects/dedal/.

Systems Biology

Prenatal Bisphenol A Exposure in Mice Induces Multi-tissue Multi-omics Disruptions Linking to Cardiometabolic Disorders

The health impacts of endocrine disrupting chemicals (EDCs) remain debated and their tissue and molecular targets are poorly understood. Here we leveraged systems biology approaches to assess the target tissues, molecular pathways, and gene regulatory networks associated with prenatal exposure to the model EDC Bisphenol A (BPA). Prenatal BPA exposure led to scores of transcriptomic and methylomic alterations in the adipose, hypothalamus, and liver tissues in mouse offspring, with cross-tissue perturbations in lipid metabolism as well as tissue-specific alterations in histone subunits, glucose metabolism and extracellular matrix. Network modeling prioritized main molecular targets of BPA, including Pparg, Hnf4a, Esr1, and Fasn. Lastly, integrative analyses identified the association of BPA molecular signatures with cardiometabolic phenotypes in mouse and human. Our multi-tissue, multi-omics investigation provides strong evidence that BPA perturbs diverse molecular networks in central and peripheral tissues, and offers insights into the molecular targets that link BPA to human cardiometabolic disorders.\n\nAuthor summaryThe inability to pinpoint the mechanistic underpinnings of environmentally-induced diseases likely stems from the pleiotropic effects of chemicals such as BPA on diverse tissues and molecular space (transcriptome, epigenome, etc.). This makes it challenging to fully dissect their health impact and merits a call for modern big data approaches to examine environmental factors. Our data-driven study is the first unbiased, multi-tissue multiomic systems biology investigation of the molecular circuitry and mechanisms underlying offspring response to prenatal BPA exposure. Importantly, the incorporation of network-based modeling allows us to capture novel players in the regulation of BPA activities in vivo, and the integration with human disease association datasets helps bridge the molecular pathways affected by BPA with diverse human diseases. In doing so, our study provides compelling molecular evidence that developmental BPA exposure significantly perturbs metabolic and endocrine systems in the offspring, and supports BPA as one of the environmental factors involved in the developmental origins of health and disease (DOHaD).

systems biology

Classification of breast tumours into molecular apocrine, luminal and basal groups based on an explicit biological model

The gene expression profiles of human breast tumours fall into three main groups that have been called luminal, basal and either HER2-enriched or molecular apocrine. To escape from the circularity of descriptive classifications based purely on gene signatures I describe a biological classification based on a model of the mammary lineage. In this model I propose that the third group is a tumour derived from a mammary hormone-sensing cell that has undergone apocrine metaplasia. I first split tumours into hormone sensing and milk secreting cells based on the expression of transcription factors linked to cell identity (the luminal progenitor split), then split the hormone sensing group into luminal and apocrine groups based on oestrogen receptor activity (the luminal-apocrine split). I show that the luminal-apocrine-basal (LAB) approach can be applied to microarray data (186 tumours) from an EORTC trial and to RNA-seq data from TCGA (674 tumours), and compare results obtained with the LAB and PAM50 approaches. Unlike pure signature-based approaches, classification based on an explicit biological model has the advantage that it is both refutable and capable of meaningful improvement as biological understanding of mammary tumorigenesis improves.

cancer biology

Analysis of the impact of molecular motions on the efficiency of XL-MS and the distance restraints in hybrid structural biology

Covalent cross-link mapping by mass spectrometry (XL-MS) is rapidly becoming the most widely used method of hybrid structural biology. We investigated the impact of incremental variations of cross-linker length have on the depth of XL-MS interrogation of protein-protein complexes, and assessed the role molecular motions in solution play in generation of cross-link-derived distance restraints. Supplementation of a popular NHS-ester cross-linker, DSS, with 2 reagents shorter or longer by CH2-CH2, increased the number of non-reductant cross-links by ~50%. Molecular dynamics simulations of these cross-linkers revealed 3 individual, partially overlapping ranges of motion, consistent with partially overlapping sets of cross-links formed by each reagent. Similar simulations elucidated protein fold-specific ranges of motions for the reactive and backbone atoms from rigid and flexible target domains. Together these findings create a quantitative framework for generation of cross-linker- and protein fold-specific distance restraints for XL-MS-guided protein-protein docking.

molecular biology

Information Processing by Simple Molecular Motifs and Susceptibility to Noise

Biological organisms rely on their ability to sense and respond appropriately to their environment. The molecular mechanisms that facilitate these essential processes are however subject to a range of random effects and stochastic processes, which jointly affect the reliability of information transmission between receptors and e.g. the physiological downstream response. Information is mathematically defined in terms of the entropy; and the extent of information flowing across an information channel or signalling system is typically measured by the \"mutual information\", or the reduction in the uncertainty about the output once the input signal is known. Here we quantify how extrinsic and intrinsic noise affect the transmission of simple signals along simple motifs of molecular interaction networks. Even for very simple systems the effects of the different sources of variability alone and in combination can give rise to bewildering complexity. In particular extrinsic variability is apt to generate \"apparent\" information that can in extreme cases mask the actual information that for a single system would flow between the different molecular components making up cellular signalling pathways. We show how this artificial inflation in apparent information arises and how the effects of different types of noise alone and in combination can be understood.

Systems Biology

Free energy calculations of protein-water complexes with Gromacs

We used GFP(Green Fluorescent Protein) to understand its basic structure, adding solvent water around the GFP, minimize and equilibrating it using molecular dynamics simulation with Gromacs. Gromacs is an open source software and widely used in molecular dynamics simulation of biological molecules such as proteins, and nucleic acids (DNA AND RNA-molecules). We employ the CHARMM (Chemistry at HARvard Molecular Mechanics) program for the force fields which enable the potential energy of a molecular system to be calculated rapidly. In this particular simple molecular mechanics interaction fields (CHARMM), the force fields consists of stretching energy, bending energy and torsion energy. The non-bond interaction energy is modeled by Lennard-Jones potential. The pdb2gmx gromacs command is implemented to obtained the basic coordinate file and topology for the particular system from the GFP PDB file (1gfl.pdb). We enclosed water molecule in a rhombic dodecahedron box having size 0.5nm, and protein are embedded in solvent water. The steepest descent method (first-order minimization) are implemented to calculate the local energy minimum with 104-steps. The stability of protein with solvent water molecule are analyzed by measuring the root-mean square displacement (RMSD) of all atoms. It was shown with help of figure that RMSD initially increases rapidly in the first part of simulation, but become stable around 0.14nm, roughly the resolution of the X-ray structure. The difference is partly due to the moving and vibration of atoms around an equilibrium structure. Secondary structure of GFP protein is presented with the help of DSSP-program.

biophysics

Twelve Elements of Visualization and Analysis for Tertiary and Quaternary Structure of Biological Molecules

During the last decades, 3D Molecular Graphics in Life Sciences has been used almost exclusively by experts through complex software and applications ranging from Structural Biology to Computer Aided Drug Design. The emergence of JavaScript and WebGL as a viable platform has enabled 3D visualization of biomolecular structures through Web browsers, without any need for specialized software. Although still in its infancy, Web Molecular Graphics opens new perspectives. This white paper, proposes a set of Twelve Elements to consider to enable 3D visualization and structural analyses of biological systems in Web molecular viewers. The Elements go beyond 3D graphics and propose an integrated approach to visualize and analyze molecular entities and their interactions in multiple dimensions, at multiple levels of details, for diverse users. The bridging of 1D sequence browsers and 3D structure viewers, possible under a Web browser, enables information flow where molecular biologists can use structural information directly at the sequence level. Given the tsunami of sequence information linked to diseases from next generation sequencing - in need for interpretation - making structural information readily available to research scientists is a tremendous opportunity for medical discovery. The Twelve Elements are conceptual and are intended to entice developers to architect software components and APIs, and to gather together as a community around common goals and open source software. A few features of emerging viewers, all available as open source, are highlighted. Speed and quality of 3D graphics for large molecular systems, the interoperability of Web components, and the instantaneous sharing of annotated visualizations through the Web, are some of the most amazing and promising capabilities of 3D Web viewing, opening bright perspectives for Life Sciences research.

bioinformatics