Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Molecular Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,693 records · Page 94Linked to original sources

Classifying Bladder Cancer Subtypes

Urothelial carcinoma of the bladder is is estimated to have killed over 16,000 people in the United States in 2016. Like breast cancer, bladder cancer is a heterogeneous disease, and characterization of its various subtypes can be useful for forecasting prognosis and treatment efficacy. According to The Cancer Genome Atlas (TCGA) project, the mRNA expression profiles of bladder tumours can be used to cluster the tumors into four different categories: I - Papillary-like, II -Luminal A, III - basal/squamous-like, and IV - other (similar to III). However it is not clear whether these mRNA expression based clusters correlate with other molecular and genetic features of the tumor cells. In other words, do differences in mRNA expression profile contain the same information as differences in protein expression, micro RNA (miRNA) expression, copy number variation and somatic mutation data. We tried to recreate mRNA based bladder tumor clusters from other multi-omic data for 328 bladder cancer tumor samples using a special deep and wide belief network composed of restricted Boltzmann machines and a multilayer perceptron. For 10-fold cross validation, we got 79% average test accuracy which implies that that differences in mRNA expression between bladder tumor cells can be reliably, though not perfectly, inferred from different molecular and genetic features of the tumors.

cancer biology

Delving deeper: Relating the behaviour of a metabolic system to the properties of its components using symbolic metabolic control analysis

High-level behaviour of metabolic systems results from the properties of, and interactions between, numerous molecular components. Reaching a complete understanding of metabolic behaviour based on the systems components is therefore a difficult task. This problem can be tackled by constructing and subsequently analysing kinetic models of metabolic pathways since such models aim to capture all the relevant properties of the system components and their interactions.\n\nSymbolic control analysis is a framework for analysing pathway models in order to reach a mechanistic understanding of their behaviour. By providing algebraic expressions for the sensitivities of system properties, such as metabolic fluxor steady-state concentrations, in terms of the properties of individual reactions it allows one to trace the high level behaviour back to these low level components. Here we apply this method to a model of pyruvate branch metabolism in Lactococcus lactis in order to explain a previously observed negative flux response towards an increase in substrate concentration. With this method we are able to show, first, that the sensitivity of flux towards changes in reaction rates (represented by flux control coefficients) is determined by the individual metabolic branches of the pathway, and second, how the sensitivities of individual reaction rates towards their substrates (represented by elasticity coefficients) contribute to this flux control. We also quantify the contributions of enzyme binding and mass-action to enzyme elasticity separately, which allows for an even finer-grained understanding of flux control.\n\nThese analytical tools allow us to analyse the control properties of a metabolic model and to arrive at a mechanistic understanding of the quantitative contributions of each of the enzymes to this control. Our analysis provides an example of the descriptive power of the general principles of symbolic control analysis.\n\nAuthor summaryMetabolic networks are complex systems consisting of numerous individual molecular components. The properties of these components, together with their non-linear interactions, give rise to high-level observed behaviour of the system in which they reside. Therefore, in order to fully understand the behaviour of a metabolic system, one has to consider the properties of all of its components. The analysis of computer models that capture and represent these systems and their components simplifies this task by allowing for an easy way to isolate the effects of each individual component. In this paper we use the framework of symbolic control analysis to investigate the sensitivity of the rate of flow of matter through one of the branches in a particular metabolic pathway towards changes in the rates of individual reactions. Here we are able to quantify how certain chains of reactions, individual reactions, and even thermodynamic and kinetic aspects of individual reactions contribute to the overall sensitivity of the rate of matter-flow. Thus, we are able to trace the behaviour of the system as a whole in a mechanistic way to the properties of the individual molecular components.

systems biology

The perils of intralocus recombination for inferences of molecular convergence

Accurate inferences of convergence require that the appropriate tree topology be used. If there is a mismatch between the tree a trait has evolved along and the tree used for analysis, then false inferences of convergence (\"hemiplasy\") can occur. To avoid problems of hemiplasy when there are high levels of gene tree discordance with the species tree, researchers have begun to construct tree topologies from individual loci. However, due to intralocus recombination even locus-specific trees may contain multiple topologies within them. This implies that the use of individual tree topologies discordant with the species tree can still lead to incorrect inferences about molecular convergence. Here we examine the frequency with which single exons and single protein-coding genes contain multiple underlying tree topologies, in primates and Drosophila, and quantify the effects of hemiplasy when using trees inferred from individual loci. In both clades we find that there are most often multiple diagnosable topologies within single exons and whole genes, with 91% of Drosophila protein-coding genes containing multiple topologies. Because of this underlying topological heterogeneity, even using trees inferred from individual protein-coding genes results in 25% and 38% of substitutions falsely labeled as convergent in primates and Drosophila, respectively. While constructing local trees can reduce the problem of hemiplasy, our results suggest that it will be difficult to completely avoid false inferences of convergence. We conclude by suggesting several ways forward in the analysis of convergent evolution, for both molecular and morphological characters.

evolutionary biology

Modular co-option of cardiopharyngeal genes during non-embryonic myogenesis

BackgroundIn chordates cardiac and body muscles arise from different embryonic origins. Myogenesis can in addition be triggered in adult organisms, during asexual development or regeneration. In the non-vertebrate ascidians, muscles originate from embryonic precursors regulated by a conserved set of genes that orchestrate cell behavior and dynamics during development. In colonial ascidians, besides embryogenesis and metamorphosis, an adult can propagate asexually via blastogenesis, skipping embryo and larval stages, and form anew the adult body, including the complete body musculature.\n\nResultsTo investigate the cellular origin and mechanisms that trigger non-embryonic myogenesis, we followed the expression of ascidian myogenic genes during Botryllus schlosseri blastogenesis, and reconstructed the dynamics of muscle precursors. Based on the expression dynamics of Tbx1/10, Ebf, Mrf, Myh3 for body wall and of FoxF, Tbx1/10, Nk4, Myh2 for heart development we show that the embryonic factors regulating myogenesis are only partially co-opted in blastogenesis, and propose that the cellular precursors contributing to heart or body muscles have different origins.\n\nConclusionsRegardless of the developmental pathway, non-embryonic myogenesis shares a similar molecular and anatomical setup as embryonic myogenesis, but implements co-option and loss of molecular modules.

developmental biology

Codon-dependent noise dictates cell-to-cell variability in nutrient poor environments

Under nutrient-rich conditions, stochasticity of transcription drives protein expression noise. However, by shifting the environment to amino acid-limited conditions, we identified in E. coli a source of noise whose strength is dictated by translational processes. Specifically, we discovered that cell-to-cell variations in fluorescent protein expression depend on codon choice, with codons yielding lower mean expression after amino acid downshift also resulting in greater noise. We propose that ultra-sensitivity in the tRNA charging/discharging cycle shapes the strength of the observed noise by amplifying fluctuations in global intracellular parameters, such as the concentrations of amino acid, synthetase, and tRNA. We hypothesize that this codon-dependent noise may allow bacteria to selectively optimize cell-to-cell variability in poor environments without relying on low molecular numbers.

systems biology

An EDS1-SAG101 complex functions in TNL-mediated immunity in Solanaceae

EDS1 (Enhanced disease susceptibility 1) forms mutually exclusive heterodimers with its interaction partners PAD4 (Phytoalexin-deficient 4) and SAG101 (Sensecence-associated gene 101). Collectively, these complexes are required for resistance responses mediated by nucleotide-binding leucine-rich repeat-type immune receptors (NLRs) possessing an N-terminal Toll-interleukin-1 receptor-like domain (TNLs). Here, immune functions of EDS1 complexes were comparatively analyzed in a mixed species approach relying on Nicotiana benthamiana (Nb), Solanum lycopersicum (Sl) and Arabidopsis thaliana (At). Genomes of most Solanaceae plants including Nb and Sl encode for two SAG101 isoforms, which engage into distinct complexes with EDS1. By a combination of genome editing and transient complementation, we show that one of these EDS1-SAG101 complexes, and not an EDS1-PAD4 complex as previously described in At, is necessary and sufficient for all tested TNL-mediated immune responses in Nb. Intriguingly, not this EDS1-SAG101 module, but mainly Solanaceae EDS1-PAD4 execute immune functions when transferred to At, and TNL functions are not restored in Nb mutant lines by expression of At EDS1 complexes. We conclude that EDS1 complexes do not represent a complete functional module, but co-evolve with additional factors, most likely protein interaction partners, for their function in TNL signaling networks of individual species. In agreement, we identify a large surface on SlEDS1 complexes required for immune activities, which may function in partner recruitment. We highlight important differences in TNL signaling networks between At and Nb, and genetic resources in the Nb system will be instrumental for future elucidation of EDS1 molecular functions.

plant biology

Self-organization controls expression more than abundance of molecular components of transcription and translation in confined cell-free gene expression

Cell-free gene expression using purified components or cell extracts has become an important platform for synthetic biology that is finding a growing numBer of practical applications. Unfortunately, at cell-relevant reactor volumes, cell-free expression suffers from excessive variability (noise) such that protein concentrations may vary by more than an order of magnitude across a population of identically constructed reaction chambers. Consensus opinion holds that variability in expression is due to the stochastic distribution of expression resources (DNA, RNAP, ribosomes, etc.) across the population of reaction chambers. In contrast, here we find that chamber-to-chamber variation in the expression efficiency generates the large variability in protein production. Through analysis and modeling, we show that chambers self-organize into expression centers that control expression efficiency. Chambers that organize into many centers, each having relatively few expression resources, exhibit high expression efficiency. Conversely, chambers that organize into just a few centers where each center has an abundance of resources, exhibit low expression efficiency. A particularly surprising finding is that diluting expression resources reduces the chamber-to-chamber variation in protein production. Chambers with dilute pools of expression resources exhibit higher expression efficiency and lower expression noise than those with more concentrated expression resources. In addition to demonstrating the means to tune expression noise, these results demonstrate that in cell-free systems, self-organization may exert even more influence over expression than the abundance of the molecular components of transcription and translation. These observations in cell-free platform may elucidate how self-organized, membrane-less structures emerge and function in cells.

biophysics

DNA methylation differences at regulatory elements are associated with the cancer risk factor age in normal breast tissue

BackgroundThe underlying biological mechanisms through which epidemiologically defined breast cancer risk factors contribute to disease risk remain poorly understood. Identification of the molecular changes associated with cancer risk factors in normal tissues may aid in determining the earliest events of carcinogenesis and informing cancer prevention strategies.\n\nResultsHere we investigated the impact cancer risk factors have on the normal breast epigenome by analyzing DNA methylation genome-wide (Infinium 450K array) in cancer-free women from the Susan G. Komen Tissue Bank (n = 100). We tested the relation of established breast cancer risk factors: age, body mass index, parity, and family history of disease with DNA methylation adjusting for potential variation in cell-type proportions. We identified 787 CpG sites that demonstrated significant associations (Q-value < 0.01) with subject age. Notably, DNA methylation was not strongly associated with the other evaluated breast cancer risk factors. Age-related DNA methylation changes are primarily increases in methylation enriched at breast epithelial cell enhancer regions (P = 7.1E-20), and binding sites of chromatin remodelers (MYC and CTCF). We validated the age-related associations in two independent populations of normal breast tissue (n = 18) and normal-adjacent to tumor tissue (n = 97). The genomic regions classified as age-related were more likely to be regions altered in cancer in both pre-invasive (n = 40, P=3.0E-03) and invasive breast tumors (n = 731, P=1.1E-13).\n\nConclusionsDNA methylation changes with age occur at regulatory regions, and are further exacerbated in cancer suggesting that age influences breast cancer risk in part through its contribution to epigenetic dysregulation in normal breast tissue.

epidemiology

A model for autonomous and non-autonomous effects of the Hippo pathway in Drosophila

While significant progress has been made toward understanding morphogen-mediated patterning in development from both the experimental and the theoretical side, the control of size and shape of tissues and organs is poorly understood. Both involve adjustment of the scale of gene expression to the size of the system, but how growth and patterning are coupled to produce scale invariance and how molecular-level information is translated into organ- and organism-level functioning is one of the most difficult problems in biology. The Hippo pathway, which controls cell proliferation and apoptosis in Drosophila and mammalian cells, contains a core kinase mechanism that affects control of the cell cycle and growth. Studies involving over- and under-expression of components in the morphogen and Hippo pathways in Drosophila reveal conditions that lead to over- or undergrowth. Herein we develop a mathematical model that incorporates the current understanding of the Hippo signal transduction network and which can explain qualitatively both the observations on whole-disc manipulations and the results arising from mutant clones. We find that a number of non-intuitive experimental results can be explained by subtle changes in the balances between inputs to the Hippo pathway. Since signal transduction and growth control pathways are highly conserved across species, much of what is learned about Drosophila applies in higher organisms, and may have direct relevance to tumor dynamics in mammalian systems.

biophysics

Pathogenic DDX3X mutations impair RNA metabolism and neurogenesis during fetal cortical development

De novo germline mutations in the RNA helicase DDX3X account for 1-3% of unexplained intellectual disability (ID) cases in females, and are associated with autism, brain malformations, and epilepsy. Yet, the developmental and molecular mechanisms by which DDX3X mutations impair brain function are unknown. Here we use human and mouse genetics, and cell biological and biochemical approaches to elucidate mechanisms by which pathogenic DDX3X variants disrupt brain development. We report the largest clinical cohort to date with DDX3X mutations (n=78), demonstrating a striking correlation between recurrent dominant missense mutations, polymicrogyria, and the most severe clinical outcomes. We show that Ddx3x controls cortical development by regulating neuronal generation and migration. Severe DDX3X missense mutations profoundly disrupt RNA helicase activity and induce ectopic RNA-protein granules and aberrant translation in neural progenitors and neurons. Together, our study demonstrates novel mechanisms underlying DDX3X syndrome, and highlights roles for RNA-protein aggregates in the pathogenesis of neurodevelopmental disease.

genetics

Proteomic Analysis of Rat Serum Revealed the Effects of Chronic Sleep Deprivation on Metabolic, Cardiovascular and Nervous System

Sleep is an essential and fundamental physiological process that plays crucial roles in the balance of psychological and physical health. Sleep disorder may lead to adverse health outcomes. The effects of sleep deprivation were extensively studied, but its mechanism is still not fully understood. The present study aimed to identify the alterations of serum proteins associated with chronic sleep deprivation, and to seek for potential biomarkers of sleep disorder mediated diseases. A label-free quantitative proteomics technology was used to survey the global changes of serum proteins between normal rats and chronic sleep deprivation rats. A total of 309 proteins were detected in the serum samples and among them, 117 proteins showed more than 1.8-folds abundance alterations between the two groups. Functional enrichment and network analyses of the differential proteins revealed a close relationship between chronic sleep deprivation and several biological processes including energy metabolism, cardiovascular function and nervous function. And four proteins including pyruvate kinase M1, clusterin, kininogen1 and profilin-1were identified as potential biomarkers for chronic sleep deprivation. The four candidates were validated via parallel reaction monitoring (PRM) based targeted proteomics. In addition, protein expression alteration of the four proteins was confirmed in myocardium and brain of rat model. In summary, the comprehensive proteomic study revealed the biological impacts of chronic sleep deprivation and discovered several potential biomarkers. This study provides further insight into the pathological and molecular mechanisms underlying sleep disorders at protein level.

bioinformatics

Lipid interactions enhance activation and potentiation of cystic fibrosis transmembrane conductance regulator (CFTR)

The recent cryo-electron microscopy structures of phosphorylated, ATP-bound CFTR in detergent micelles failed to reveal an open anion conduction pathway as expected on the basis of previous functional studies in biological membranes. We tested the hypothesis that interaction of CFTR with lipids is important for opening of its channel. Interestingly, molecular dynamics studies revealed that phospholipids associate with regions of CFTR proposed to contribute to its channel activity. More directly, we found that CFTR purified together with associated lipids using the amphipol: A8-35, exhibited higher rates of catalytic activity, channel activation and potentiation using ivacaftor, than did CFTR purified in detergent. Catalytic activity in CFTR detergent micelles was partially rescued by addition of phospholipids plus cholesterol, arguing that these lipids contribute directly to its modulation. In summary, these studies highlight the importance of lipids in regulated CFTR channel activation and potentiation.

biophysics

Human 5′ UTR design and variant effect prediction from a massively parallel translation assay

Predicting the impact of cis-regulatory sequence on gene expression is a foundational challenge for biology. We combine polysome profiling of hundreds of thousands of randomized 5' UTRs with deep learning to build a predictive model that relates human 5' UTR sequence to translation. Together with a genetic algorithm, we use the model to engineer new 5' UTRs that accurately target specified levels of ribosome loading, providing the ability to tune sequences for optimal protein expression. We show that the same approach can be extended to chemically modified RNA, an important feature for applications in mRNA therapeutics and synthetic biology. We test 35,000 truncated human 5' UTRs and 3,577 naturally-occurring variants and show that the model accurately predicts ribosome loading of these sequences. Finally, we provide evidence of 47 SNVs associated with human diseases that cause a significant change in ribosome loading and thus a plausible molecular basis for disease.

synthetic biology

Accurate estimation of molecular counts in droplet-based single-cell RNA-seq experiments

Single-cell RNA-seq protocols provide powerful means for examining the gamut of cell types and transcriptional states that comprise complex biological tissues. Recently-developed approaches based on droplet microfluidics, such as inDrop or Drop-seq, use massively multiplexed barcoding to enable simultaneous measurements of transcriptomes for thousands of individual cells. The increasing complexity of such data also creates challenges for subsequent computational processing and troubleshooting of these experiments, with few software options currently available. Here we describe a flexible pipeline for processing droplet-based transcriptome data that implements barcode corrections, classification of cell quality, and diagnostic information about the droplet libraries. We introduce advanced methods for correcting composition bias and sequencing errors affecting cellular and molecular barcodes to provide more accurate estimates of molecular counts in individual cells.

genomics

A prior-based approach for hypothesis comparison and its utility to discern among temporal scenarios of divergence

One of the major problems in evolutionary biology is to elucidate the relationships between historical events and the tempo and mode of lineage divergence. The development of relaxed molecular clock models and the increasing availability of DNA sequences resulted in more accurate estimations of taxa divergence times. However, finding the link between competing historical events and divergence is still challenging. Here we investigate assigning constrained-age priors to nodes of interest in a time-calibrated phylogeny as a means of hypothesis comparison. These priors are equivalent to historic scenarios for lineage origin. The hypothesis that best explains the data can be selected by comparing the likelihood values of the competing hypotheses, modelled with different priors. A simulation approach was taken to evaluate the performance of the prior-based method and to compare it with an unconstrained approach. We explored the effect of DNA sequence length and the temporal placement and span of competing hypotheses (i.e. historic scenarios) on selection of the correct hypothesis and the strength of the inference. Competing hypotheses were compared applying a posterior simulation analogue of the Akaike Information Criterion and Bayes factors (obtained after calculation of the marginal likelihood with three estimators: Harmonic Mean, Stepping Stone and Path Sampling). We illustrate the potential application of the prior-based method on an empirical data set to compare competing geological hypotheses explaining the biogeographic patterns in Pleurodeles newts. The correct hypothesis was selected on average 89% times. The best performance was observed with DNA sequence length of 3500-10000 bp. The prior-based method is most reliable when the hypotheses compared are not temporally too close. The strongest inferences were obtained when using the Stepping Stone and Path Sampling estimators. The prior-based approach proved effective in discriminating between competing hypotheses when used on empirical data. The unconstrained analyses performed well but it probably requires additional computational effort. Researchers applying this approach should rely only on inferences with moderate to strong support. The prior-based approach could be applied on biogeographical and phylogeographical studies where robust methods for historical inferences are still lacking.

evolutionary biology

A molecular framework for functional versatility of HECATE transcription factors

During the plant life cycle, diverse signalling inputs are continuously integrated and engage specific genetic programs depending on the cellular or developmental context. Consistent with an important role in this process, HECATE (HEC) bHLH transcription factors display diverse functions, from photomorphogenesis to the control of shoot meristem dynamics and gynoecium patterning. However, the molecular mechanisms underlying their functional versatility and the deployment of specific HEC sub-programs still remain elusive.\n\nTo address this issue, we systematically identified proteins with the capacity to interact with HEC1, the best characterized member of the family, and integrated this information with our data set of direct HEC1 target genes. The resulting core genetic modules were consistent with specific developmental functions of HEC1, including its described activities in light signalling, gynoecium development and auxin homeostasis. Importantly, we found that in addition, HEC genes play a role in the modulation of flowering time and uncovered that their role in gynoecium development may involve the direct transcriptional regulation of NGATHA1 (NGA1) and NGA2 genes. NGA factors were previously shown to contribute to fruit development, but our data now show that they also modulate stem cell homeostasis in the SAM.\n\nTaken together, our results suggest a molecular network underlying the functional versatility of HEC transcription factors. Our analyses have not only allowed us to identify relevant target genes controlling shoot stem cell activity and a so far undescribed biological function of HEC1, but also provide a rich resource for the mechanistic elucidation of further context dependent HEC activities.\n\nSignificance statementAlthough many transcription factors display diverse regulatory functions during plant development, our understanding of the underlying mechanisms remains poor. Here, by reconstructing the regulatory modules orchestrated by the bHLH transcription factor HECATE1 (HEC1), we defined its regulatory signatures and delineated a regulatory network that provides a molecular basis for its functional versatility. In addition, we uncovered a function for HEC genes in modulating flowering time and further identified downstream signalling components balancing shoot stem cell activity.

plant biology

Integrated analysis of single-cell embryo data yields a unified transcriptome signature for the human preimplantation epiblast

Single-cell profiling techniques create opportunities to delineate cell fate progression in mammalian development. Recent studies provide transcriptome data from human preimplantation embryos, in total comprising nearly 2000 individual cells. Interpretation of these data is confounded by biological factors such as variable embryo staging and cell-type ambiguity, as well as technical challenges in the collective analysis of datasets produced with different sample preparation and sequencing protocols. Here we address these issues to assemble a complete gene expression time course spanning human preimplantation embryogenesis. We identify key transcriptional features over developmental time and elucidate lineage-specific regulatory networks. We resolve post hoc cell-type assignment in the blastocyst, and define robust transcriptional prototypes that capture epiblast and primitive endoderm lineages. Examination of human pluripotent stem cell transcriptomes in this framework identifies culture conditions that sustain a naive state pertaining to the inner cell mass. Our approach thus clarifies understanding both of lineage segregation in the early human embryo and of in vitro stem cell identity, and provides an analytical resource for comparative molecular embryology.

developmental biology

The genetics of resistance to Morinda fruit toxin during the postembryonic stages in Drosophila sechellia

Many phytophagous insect species are ecologic specialists that have adapted to utilize a single host plant. Drosophila sechellia is a specialist that utilizes the ripe fruit of Morinda citrifolia, which is toxic to its sibling species, D. simulans. Here we apply multiplexed shotgun genotyping and QTL analysis to examine the genetic basis of resistance to M. citrifolia fruit toxin in interspecific hybrids. We find that at least four dominant and four recessive loci interact additively to confer resistance to the M. citrifolia fruit toxin. These QTL include a dominant locus of large effect on the third chromosome (QTL-IIIsima) that was not detected in previous analyses. The small-effect loci that we identify overlap with regions that were identified in selection experiments with D. simulans on octanoic acid and in QTL analyses of adult resistance to octanoic acid. Our high-resolution analysis sheds new light upon the complexity of M. citrifolia resistance, and suggests that partial resistance to lower levels of M. citrifolia toxin could be passed through introgression from D. sechellia to D. simulans in nature. The identification of a locus of major effect, QTL-IIIsima, is an important step towards identifying the molecular basis of host plant specialization by D. sechellia.

Evolutionary Biology