Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Molecular Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,333 records · Page 74Linked to original sources

Conceptual Confusion: the case of Epigenetics

The observations of phenotypic plasticity have stimulated the revival of epigenetics. Over the past 70 years the term has come in many colors and flavors, depending on the biological discipline and time period. The meanings span from Waddingtons \"epigenotype\" and \"epigenetic landscape\" to the molecular biologists \"epigenetic marks\" embodied by DNA methylation and histone modifications. Here we seek to quell the ambiguity of the name. First we place \"epigenetics\" in the various historical contexts. Then, by presenting the formal concepts of dynamical systems theory we show that the \"epigenetic landscape\" is more than a metaphor: it has specific mathematical foundations. The latter explains how gene regulatory networks produce multiple attractor states, the self-stabilizing patterns of gene activation across the genome that account for \"epigenetic memory\". This network dynamics approach replaces the reductionist correspondence of molecular epigenetic modifications with concept of the epigenetic landscape, by providing a concrete and crisp correspondence.

Systems Biology

The impact of tumor receptor heterogeneity on the response to anti-angiogenic cancer treatment

Multiple promoters and inhibitors mediate angiogenesis, the formation of new blood vessels, and these factors represent potential targets for impeding vessel growth in tumors. Vascular endothelial growth factor (VEGF) is a potent angiogenic factor targeted in anti-angiogenic cancer therapies. In addition, thrombospondin-1 (TSP1) is a major endogenous inhibitor of angiogenesis, and TSP1 mimetics are being developed as an alternative type of anti-angiogenic agent. The combination of bevacizumab, an anti-VEGF agent, and ABT-510, a TSP1 mimetic, has been tested in clinical trials to treat advanced solid tumors. However, the patients responses are highly variable and show disappointing outcomes. To obtain mechanistic insight into the effects of this combination anti-angiogenic therapy, we have constructed a novel whole-body systems biology model including the VEGF and TSP1 reaction networks. Using this molecular-detailed model, we investigated how the combination anti-angiogenic therapy changes the amounts of pro-angiogenic and anti-angiogenic complexes in cancer patients. We particularly focus on answering the question of how the effect of the combination therapy is influenced by tumor receptor expression, one aspect of patient-to-patient variability. Overall, this model complements the clinical administration of combination anti-angiogenic therapy, highlights the role of tumor receptor variability in the heterogeneous responses to anti-angiogenic therapy, and identifies the tumor receptor profiles that correlate with a high likelihood of a positive response to the combination therapy. Our model provides novel understanding of the VEGF-TSP1 balance in cancer patients at the systems-level and could be further used to optimize combination anti-angiogenic therapy.

systems biology

Improving the consistency of functional genomics screens using molecular features - a multi-omics, pan-cancer study

Probing the genetic dependencies of cancer cells helps understand the tumor biology and identify potential drug targets. RNAi-based shRNA and CRISPR/Cas9-based sgRNA have been commonly utilized in functional genetic screens to identify essential genes affecting growth rates in cancer cell lines. However, questions remain whether the gene essentiality profiles determined using these two technologies are comparable. In the present study, we collected 42 cell lines representing a variety of 10 tissue types, which had been screened both by shRNA and CRISPR techniques. We observed poor consistency of the essentiality scores between the two screens for the majority of the cell lines. The consistency did not improve after correcting the off-target effects in the shRNA screening, suggesting a minimal impact of off-target effects. We considered a linear regression model where the shRNA essentiality score is the predictor and the CRISPR essentiality score is the response variable. We showed that by including molecular features such as mutation, gene expression and copy number variation as covariates, the predictability of the regression model greatly improved, suggesting that molecular features may provide critical information in explaining the discrepancy between the shRNA and CRISPR-based essentiality scores. We provided a Combined Essentiality Score (CES) based on the model prediction and showed that the CES greatly improved the consensus of common essential genes. Furthermore, the CES also identified novel essential genes that are specific to individual cell types. Taken together, we provided a systematic approach to define a more accurate gene essentiality profile by integrating functional screen data and molecular profiles.

systems biology

Promiscuity of peripheral molecular biomarkers in major psychiatric disorders: a transdiagnostic systematic review

The search for biomarkers has been one of the leading endeavors in biological psychiatry; nevertheless, in spite of hundreds of publications, it has yet to make an impact in clinical practice. To study how biomarker research has progressed over the years, we performed a systematic review of the literature to evaluate (a) the most studied peripheral molecular markers in major psychiatric disorders, (b) the main features of studies in which they are proposed as biomarkers and (c) whether their patterns of variation are similar across disorders. Out of the six molecules most commonly present as keywords in articles studying plasmatic markers of schizophrenia, major depressive disorder or bipolar disorder, five (BDNF, TNF-alpha, IL-6, C-reactive protein and cortisol) were the same across the three diagnoses. An analysis of the literature on these molecules showed that, while 65% of studies were cross-sectional and 66% compared biomarker levels between patients and controls in specific disorders, only 10% presented an objective measure of diagnostic or prognostic efficacy. Meta-analyses showed that variation in the levels of these molecules was robust across studies, but also similar among disorders, suggesting them to reflect transdiagnostic systemic consequences of psychiatric illness rather than diagnostic markers. Based on this, we discuss how current publication practices have led to research fragmentation across diagnoses, and what steps can be taken in order to increase clinical translation in the field.

neuroscience

High-Performance Image-Based Measurements of Biological Forces and Interactions in a Dual Optical Trap

Optical traps enable nanoscale manipulation of individual biomolecules while measuring molecular forces and lengths. This ability relies on the sensitive detection of optically trapped particles, typically accomplished using laser-based interferometric methods. Recently, precise and fast image-based particle tracking techniques have garnered increased interest as a potential alternative to laser-based detection, however successful integration of image-based methods into optical trapping instruments for biophysical applications and force measurements has remained elusive. Here we develop a camera-based detection platform that enables exceptionally accurate and precise measurements of biological forces and interactions in a dual optical trap. In demonstration, we stretch and unzip DNA molecules while measuring the relative distances of trapped particles from their trapping centers with sub-nanometer accuracy and precision, a performance level previously only achieved using photodiodes. We then use the DNA unzipping technique to localize bound proteins with extraordinary sub-base-pair precision, revealing how thermal DNA fluctuations allow an unzipping fork to sense and respond to a bound protein prior to a direct encounter. This work significantly advances the capabilities of image tracking in optical traps, providing a state-of-the-art detection method that is accessible, highly flexible, and broadly compatible with diverse experimental substrates and other nanometric techniques.

biophysics

Defects of myelination are common pathophysiology in syndromic and idiopathic autism spectrum disorders

Autism spectrum disorder (ASD) affects approximately 1:68 individuals and has incalculable burdens on affected individuals, their families, and health care systems. While the genetic contributions to idiopathic ASD are heterogeneous and largely unknown, the causal mutations for syndromic forms of ASD - including truncations and copy number variants - provide a genetic toehold with which to gain mechanistic insights1-3. Models of these syndromic disorders have been used to better characterize the molecular and physiological processes disrupted by these mutations4. Two fundamental questions remain - how biologically similar are the mouse models of syndromic forms of ASD, and how relevant are these mouse models to their human analogs? To address these questions, we performed integrative transcriptomic analyses of seven independent mouse models of three syndromic forms of ASD generated across five laboratories, and assessed dysregulated genes and their pathways in human postmortem brain from patients with ASD and unaffected controls. These cross-species analyses converged on shared disruptions in myelination and axon development across both syndromic and idiopathic ASD, highlighting both the face validity of mouse models for these disorders and identifying novel convergent molecular phenotypes amendable to rescue with therapeutics.

neuroscience

Data aggregation at the level of molecular pathways improves stability of experimental transcriptomic and proteomic data

High throughput technologies opened a new era in biomedicine by enabling massive analysis of gene expression at both RNA and protein levels. Unfortunately, expression data obtained in different experiments are often poorly compatible, even for the same biological samples. Here, using experimental and bioinformatic investigation of major experimental platforms, we show that aggregation of gene expression data at the level of molecular pathways helps to diminish cross- and intra-platform bias otherwise clearly seen at the level of individual genes. We created a mathematical model of cumulative suppression of data variation that predicts the ideal parameters and the optimal size of a molecular pathway. We compared the abilities to aggregate experimental molecular data for the five alternative methods, also evaluated by their capacity to retain meaningful features of biological samples. The bioinformatic method OncoFinder showed optimal performance in both tests and should be very useful for future cross-platform data analyses.

Bioinformatics

Identification of biological mechanisms by semantic classifier systems

The interpretability of a classification model is one of its most essential characteristics. It allows for the generation of new hypotheses on the molecular background of a disease. However, it is questionable if more complex molecular regulations can be reconstructed from such limited sets of data. To bridge the gap between complexity and interpretability, we replace the de novo reconstruction of these processes by a hybrid classification approach partially based on existing domain knowledge. Using semantic building blocks that reflect real biological processes these models were able to construct hypotheses on the underlying genetic configuration of the analysed phenotypes. As in the building process, also these hypotheses are composed of high-level biology-based terms. The semantic information we utilise from gene ontology is a vocabulary which comprises the essential processes or components of a biological system. The constructed semantic multi-classifier system consists of expert base classifiers which each select the most suitable term for characterising their assigned problems. Our experiments conducted on datasets of three distinct research fields revealed terms with well-known associations to the analysed context. Furthermore, some of the chosen terms do not seem to be obviously related to the issue and thus lead to new, hypotheses to pursue.\n\nAuthor summaryData mining strategies are designed for an unbiased de novo analysis of large sample collections and aim at the detection of frequent patterns or relationships. Later on, the gained information can be used to characterise diagnostically relevant classes and for providing hints to the underlying mechanisms which may cause a specific phenotype or disease. However, the practical use of data mining techniques can be restricted by the available resources and might not correctly reconstruct complex relationships such as signalling pathways.\n\nTo counteract this, we devised a semantic approach to the issue: a multi-classifier system which incorporates existing biological knowledge and returns interpretable models based on these high-level semantic terms. As a novel feature, these models also allow for qualitative analysis and hypothesis generation on the molecular processes and their relationships leading to different phenotypes or diseases.

systems biology

Gene co-expression networks in whole blood implicate multiple interrelated molecular pathways in obese asthma

BackgroundAsthmatic children who develop obesity have poorer outcomes compared to those that do not, including poorer control, more severe symptoms, and greater resistance to standard treatment. Gene expression networks are powerful statistical tools for characterizing the underpinnings of human disease that leverage the putative co-regulatory relationships of genes to infer biological pathways altered in disease states.\n\nObjectiveThe aim of this study was to characterize the biology of childhood asthma complicated by adult obesity.\n\nMethodsWe performed weighted gene co-expression network analysis (WGCNA) of gene expression data in whole blood from 514 adult subjects from the Childhood Asthma Management Program (CAMP). We then performed module preservation and association replication analyses in 418 subjects from two independent asthma cohorts (one pediatric and one adult).\n\nResultsWe identified a multivariate model in which four gene co-expression network modules were associated with incident obesity in CAMP (each P < 0.05). The module memberships were enriched for genes in pathways related to platelets, integrins, extracellular matrix, smooth muscle, NF-{kappa}B signaling, and Hedgehog signaling. The network structures of each of the four obese asthma modules were significantly preserved in both replication cohorts (permutation P = 9.999E-05). The corresponding module gene sets were significantly enriched for differential expression in obese subjects in both replication cohorts (each P < 0.05).\n\nConclusionsOur gene co-expression network profiles thus implicate multiple interrelated pathways in the biology of an important endotype of obese asthma.\n\nKey MessagesO_LIWe hypothesized that individuals with asthma complicated by obesity had distinct blood gene expression signatures.\nC_LIO_LIGene co-expression network analysis implicated several inflammatory biological pathways in one form of obese asthma.\nC_LI\n\nCapsule SummaryThis work addresses a knowledge gap about the molecular relationship between asthma and obesity, suggesting that an endotype of obese asthma, known as asthma complicated by obesity, is underpinned by coherent biological mechanisms.\n\nAbbreviations

genomics

Two biological constants for accurate classification and evolution pattern analysis of Subgen.strobus and subgen. Pinus

Currently, biological classification and determination of different categories are all based on empirical knowledge,which is obtained relying on morphological and molecular characters. For these methods they lacks of absolutely quantitative criteria ground on intrinsically scientific principles. In fact, accurate science classification must depend on the correct description of biology evolution rules.\n\nIn this article a new theoretical approach was proposed, in which two characteristic constants were gained from biological common heredity and variation information theory equation, when it is at the maximum information states, corresponding to symmetric and asymmetric variation states. They are common composition ratios, Pg =0.61, and Pg=0.70. By analyzing the common composition ratios of compounds among oleoresins, two pine subgenus:Subgen.Strobus (Sweet) Held and Subgen. Pinus could be integrated into one class, Genus pinus, excellently, when Pg = 0.61.\n\nThese two pine subgenus could be classified into two groups clearly,when Pg = 0.70.\n\nThe results is somewhat different from that achieved by means of classical classification relying on morphological characters. On the other hand, the evolution relationship of two subgenus was analyzed based on characteristic sequences of samples, it indicated that white pine origin from pinus tabuliformis. The two constants should be used as the classification constants of some biological categories of plants.

bioinformatics

Effects of duplicated mapped read PCR artifacts on RNA-seq differential expression analysis based on qRNA-seq

Best practices to handling duplicated mapped reads in RNA-seq analyses has long been discussed but a gold standard method has yet to be established, as such duplicates could originate from valid biological transcripts or they could be PCR-related artifacts. Here we used the NEXTflex qRNA-SeqTM (aka Molecular Indexing) technology to identify PCR duplicates via the random attachment of unique molecular labels to each cDNA molecule prior to PCR amplification. We found that up to 64.3% of the single end and 19.3% of the mouse paired end duplicates originated from valid biological transcripts rather than PCR artifacts. For single end reads, either removing or retaining all duplicates resulted in a substantial number of false positives (up to 47.0%) and false negatives (up to 12.1%) in the sets of significantly differentially expressed genes. For paired end reads, only the alignment retaining all duplicates resulted in a substantial number of false positives. This is the first effort to evaluate the performance of qRNA-seq using real-world biomedical samples, and we found that PCR duplicate identification provided minor benefits for paired end reads but greatly improved the sensitivity and specificity in the determination of the significantly differentially expressed genes for single end reads.

genomics

Time course analysis of the brain transcriptome during transitions between brood care and reproduction in the clonal raider ant

Division of labor between reproductive queens and non-reproductive workers that perform brood care is the hallmark of insect societies. However, the molecular basis of this fundamental dichotomy remains poorly understood, in part because the caste of an individual cannot typically be experimentally manipulated at the adult stage. Here we take advantage of the unique biology of the clonal raider ant, Ooceraea biroi, where reproduction and brood care behavior can be experimentally manipulated in adults. To study the molecular regulation of reproduction and brood care, we induced transitions between both states, and monitored brain gene expression at multiple time points. We found that introducing larvae that inhibit reproduction and induce brood care behavior caused much faster changes in adult gene expression than removing larvae. The delayed response to the removal of the larval signal prevents untimely activation of reproduction in O. biroi colonies. This resistance to change when removing a signal also prevents premature modifications in many other biological processes. Furthermore, we found that the general patterns of gene expression differ depending on whether ants transition from reproduction to brood care or vice versa, indicating that gene expression changes between phases are cyclic rather than pendular. Our analyses also identify genes with large and early expression changes in one or both transitions. These genes likely play upstream roles in regulating reproduction and behavior, and thus constitute strong candidates for future molecular studies of the evolution and regulation of reproductive division of labor in insect societies.

evolutionary biology

Sexual dimorphism in the Drosophila metabolome increases throughout development

The expression of sexually dimorphic phenotypes from a shared genome between males and females is a longstanding puzzle in evolutionary biology. Increasingly, research has made use of transcriptomic technology to examine the molecular basis of sexual dimorphism through gene expression studies, but even this level of detail misses the metabolic processes that ultimately link gene expression with the whole organism phenotype. We use metabolic profiling in Drosophila melanogaster to complete this missing step, with a view to examining variation in male and female metabolic profiles, or metabolomes, throughout development. We show that the metabolome varies considerably throughout larval, pupal and adult stages. We also find significant sexual dimorphism in the metabolome, although only in pupae and adults, and the extent of dimorphism tends to increase throughout development. We compare this to transcriptomic data from the same population and find that the general pattern of increasing sex differences throughout development is mirrored in RNA expression. We discuss our results in terms of the usefulness of metabolic profiling in linking genotype and phenotype to more fully understand the basis of sexually dimorphic phenotypes.

Physiology

Molecular adaptation in Rubisco: discriminating between convergent evolution and positive selection using mechanistic and classical codon models

Rubisco (Ribulose-1, 5-biphosphate carboxylase/oxygenase) is the most important enzyme on earth, catalyzing the first step of CO2 fixation in photosynthesis. Its molecular adaptation to C4 photosynthetic pathway has attracted a lot of attention. C4 plants, which comprise less than 5% of land plants, have evolved more efficient photosynthesis compared to C3 plants. Interestingly, a large number of independent transitions from C3 to C4 phenotype have occurred. Each time, the Rubisco enzyme has been subject to similar changes in selective pressure, thus providing an excellent model for convergent evolution at the molecular level. Molecular adaptation is often identified with positive selection and is typically characterized by an elevated ratio of non-synonymous over synonymous substitution rates (dN/dS). However, convergent adaptation is expected to leave a different molecular signature, taking the form of repeated transitions toward identical or similar amino acids.\n\nHere, we use a previously introduced codon-based differential selection model to detect and quantify consistent patterns of convergent adaptation in Rubisco in Amaranthaceae. We further contrast the results thus obtained with those obtained under classical codon models based on the estimation of dN/dS. We find that the two classes of models tend to select distinct, although overlapping, sets of positions. This discrepancy in the results illustrates the conceptual difference between these models, while emphasizing the need to better discriminate between qualitatively different selective regimes, by using a broader class of codon models than those currently considered in molecular evolutionary studies.

Evolutionary Biology

Adaptive Evolution Of Transcription Factor Binding Affinities

Understanding the molecular basis of gene expression evolution is a central problem in evolutionary biology. However, connecting changes in gene expression to increased fitness, and identifying the functional basis of those changes, remains challenging. To study adaptive evolution of gene expression in real time, we performed long term experimental evolution (LTEE) of Saccharomyces cerevisiae (budding yeast) in ammonium-limited chemostats. Following several hundred generations of continuous selection we found significant divergence of nitrogen-responsive gene expression in lineages with increased fitness. In multiple independent lineages we found repeated selection for non-synonymous mutations in the zinc finger DNA binding domain of the activating transcription factor (TF), GAT1, that operates within incoherent feedforward loops to control expression of the nitrogen catabolite repression (NCR) regulon. Missense mutations in the DNA binding domain of GAT1 reduce its binding affinity for the GATAA consensus sequence in a promoter-specific manner, resulting in increased expression of ammonium permease genes via both direct and indirect effects, thereby conferring increased fitness. We find that altered transcriptional output of the NCR regulon results in antagonistic pleiotropy in alternate environments and that the DNA binding domain of GAT1 is subject to purifying selection in natural populations. Our study shows that adaptive evolution of gene expression can entail tuning expression output by quantitative changes in TF binding affinities while maintaining the overall topology of a gene regulatory network.

evolutionary biology

RGBM: Regularized Gradient Boosting Machines For The Identification of Transcriptional Regulators Of Discrete Glioma Subtypes

The transcription factors (TF) which regulate gene expressions are key determinants of cellular phenotypes. Reconstructing large-scale genome-wide networks which capture the influence of TFs on target genes are essential for understanding and accurate modelling of living cells. We propose RGBM: a gene regulatory network (GRN) inference algorithm, which can handle data from heterogeneous information sources including dynamic time-series, gene knockout, gene knockdown, DNA microarrays and RNA-Seq expression profiles. RGBM allows to use an a priori mechanistic of active biding network consisting of TFs and corresponding target genes. RGBM is evaluated on the DREAM challenge datasets where it surpasses the winners of the competitions and other established methods for two evaluation metrics by about 10-15%.\n\nWe use RGBM to identify the main regulators of the molecular subtypes of brain tumors. Our analysis reveals the identity and corresponding biological activities of the master regulators driving transformation of the G-CIMP-high into the G-CIMP-low subtype of glioma and PA-like into LGm6-GBM, thus, providing a clue to the yet undetermined nature of the transcriptional events driving the evolution among these novel glioma subtypes.\n\nRGBM is available for download on CRAN at https://cran.rproject.org/web/packages/RGBM/index.html

bioinformatics

A genetic screen suggests an alternative mechanism for inhibition of SecA by azide

Sodium azide prevents bacterial growth by inhibiting the activity of SecA, which is required for translocation of proteins across the cytoplasmic membrane. Azide inhibits ATP turnover in vitro, but its mechanism of action in vivo is unclear. To investigate how azide inhibits SecA in cells, we used transposon directed insertion-site sequencing (TraDIS) to screen a library of transposon insertion mutants for mutations that affect the susceptibility of E. coli to azide. Insertions disrupting components of the Sec machinery generally increased susceptibility to azide, but insertions truncating the C-terminal tail (CTT) of SecA decreased susceptibility of E. coli to azide. Treatment of cells with azide caused increased aggregation of the CTT, suggesting that azide disrupts its structure. Analysis of the metal-ion content of the CTT indicated that SecA binds to iron and the azide disrupts the interaction of the CTT with iron. Azide also disrupted binding of SecA to membrane phospholipids, as did alanine substitutions in the metal-coordinating amino acids. Furthermore, treating purified phospholipid-bound SecA with azide in the absence of added nucleotide disrupted binding of SecA to phospholipids. Our results suggest that azide does not inhibit SecA by inhibiting the rate of ATP turnover in vivo. Rather, azide inhibits SecA by causing it to \"backtrack\" from the ADP-bound to the ATP-bound conformation, which disrupts the interaction of SecA with the cytoplasmic membrane.\n\nSignificance statementSecA is a bacterial ATPase that is required for the translocation of a subset of secreted proteins across the cytoplasmic membrane. Sodium azide is a well-known inhibitor of SecA, but its mechanism of action in vivo is poorly understood. To investigate this mechanism, we examined the effect of azide on the growth of a library of [~]1 million transposon insertion mutations. Our results suggest that azide causes SecA to backtrack in its ATPase cycle, which disrupts binding of SecA to the membrane and to its metal cofactor, which is iron. Our results provide insight into the molecular mechanism by which SecA drives protein translocation and how this essential biological process can be disrupted.

microbiology

Enter the matrix: Interpreting unsupervised feature learning with matrix decomposition to discover hidden knowledge in high-throughput omics data

Omics data contains signal from the molecular, physical, and kinetic inter- and intra-cellular interactions that control biological systems. Matrix factorization techniques can reveal low-dimensional structure from high-dimensional data that reflect these interactions. These techniques can uncover new biological knowledge from diverse high-throughput omics data in topics ranging from pathway discovery to time course analysis. We review exemplary applications of matrix factorization for systems-level analyses. We discuss appropriate application of these methods, their limitations, and focus on analysis of results to facilitate optimal biological interpretation. The inference of biologically relevant features with matrix factorization enables discovery from high-throughput data beyond the limits of current biological knowledge--answering questions from high-dimensional data that we have not yet thought to ask.

systems biology