Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Molecular Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,477 records · Page 82Linked to original sources

Insights on protein thermal stability: a graph representation of molecular interactions

Understanding the molecular mechanisms of thermal stability is a challenge in protein biology. Indeed, knowing the temperature at which proteins are stable has important theoretical implications, which are intimately linked with properties of the native fold, and a wide range of potential applications from drug design to the optimization of enzyme activity.\n\nHere, we present a novel graph-theoretical framework to assess thermal stability based on the structure without any a priori information. In our approach we describe proteins as energy-weighted graphs and compare them using ensembles of interaction networks. Investigating the position of specific interactions within the 3D native structure, we developed a parameter-free network descriptor that permits to distinguish thermostable and mesostable proteins with an accuracy of 76% and Area Under the Roc Curve of 78%.

bioinformatics

Transcriptional regulatory mechanisms of fibrosis development in mouse lung tissue exposed to carbon nanotubes

BackgroundCarbon nanotubes (CNTs) usage has rapidly increased in the last few decades due to their unique properties, exploited in various industrial and commercial products. Certain types of CNTs cause adverse health effects, including chronic inflammation and fibrosis. Despite the large number of in vitro and in vivo studies evaluating these effects, many important questions remain unanswered due to a lack of mechanistic understanding of how CNTs induce cellular stress responses. In order to predict CNT toxicity, it is important to understand which transcriptional programs are specifically activated in response to CNTs, and what similarities and differences exist in relation to other toxic inducers exerting similar adverse effects.\n\nResultsA systems biology approach was applied to reveal complex interactions at the molecular level in mouse lung tissue in response to different fibrosis inducers: two types of multi-walled CNTs, NM-401 and NRCWE-26, and bleomycin (BLM). Based on mRNA gene expression profiles, we inferred gene regulatory networks (GRNs) to capture functional hierarchical regulatory structures between genes and their regulators. We found that activities of the transcription factors (TFs) Myc, Arid5a and Mxd1 were associated with the regulation of cytokine transcription in response to CNTs, while in response to BLM treatment, Myc was associated with p53 signaling. TF Litaf was identified as the essential regulator for noncanonical signaling of TLR2/4 driven by CNTs. Despite the different nature of the lung injury caused by CNTs and BLM, we identified common stress response modules, that included DNA damage (TFs: E2f8, E2f1, Foxm1), M1/M2 macrophage polarization (TF: Mafb), Interferon response (TFs: Irf7, Stat2 and Irf9) for all agents.\n\nConclusionsThese results suggest that the reconstruction and analysis of TF-centric gene interaction networks can reveal key targets and regulators of cellular stress responses to toxic agents.

pharmacology and toxicology

More than meets the eye: diversity and geographic patterns in sea cucumbers

Estimates for the number of species in the sea vary by orders of magnitude. Molecular taxonomy can greatly speed up screening for diversity and evaluating species boundaries, while gaining insights into the biology of the species. DNA barcoding with a region of cytochrome oxidase 1 (COI) is now widely used as a first pass for molecular evaluation of diversity, as it has good potential for identifying cryptic species and improving our understanding of marine biodiversity. We present the results of a large scale barcoding effort for holothuroids (sea cucumbers). We sequenced 3048 individuals from numerous localities spanning the diversity of habitats in which the group occurs, with a particular focus in the shallow tropics (Indo-Pacific and Caribbean) and the Antarctic region. The number of cryptic species is much higher than currently recognized. The vast majority of sister species have allopatric distributions, with species showing genetic differentiation between ocean basins, and some are even differentiated among archipelagos. However, many closely related and sympatric forms, that exhibit distinct color patterns and/or ecology, show little differentiation in, and cannotbe separatedby, COI sequence data. This pattern is much more common among echinoderms than among molluscs or arthropods, and suggests that echinoderms acquire reproductive isolation at a much faster pace than other marine phyla. Understanding the causes behind such patterns will refine our understanding of diversification and biodiversity in the sea.

Evolutionary Biology

Gene tree discordance generates patterns of diminishing convergence over time

Phenotypic convergence is an exciting outcome of adaptive evolution, occurring when species find similar solutions to the same problem. Unraveling the molecular basis of convergence provides a way to link genotype to adaptive phenotypes, but can also shed light on the extent to which evolution is repeatable and predictable. Many recent genome-wide studies have uncovered a striking pattern of diminishing convergence over time, ascribing this pattern to the presence of intramolecular epistatic interactions. Here, we consider gene tree discordance as an alternative driver of convergence levels over time. We demonstrate that gene tree discordance can produce patterns of diminishing convergence by itself, and that controlling for discordance as a cause of apparent convergence makes the pattern disappear. We also show that synonymous substitutions, where neither selection nor epistasis should be prevalent, have the same diminishing pattern of molecular convergence among closely related primate species. Finally, we demonstrate that even in situations where biological discordance is not possible, errors in species tree inference can drive these same patterns. Though intramolecular epistasis is undoubtedly affecting many proteins, our results suggest an additional explanation for this widespread pattern. These results contribute to a growing appreciation not just of the presence of gene tree discordance, but of the unpredictable effects this discordance can have on analyses of molecular evolution.

Evolutionary Biology

Ancestral Reconstruction of Protein Interaction Networks

The molecular and cellular basis of novelty is a major open question in evolutionary biology. Until very recently, the vast majority of cellular phenomena were so difficult to sample that cross-species studies of biochemistry were rare and comparative analysis at the level of biochemical systems was almost impossible. Recent advances in systems biology are changing what is possible, however, and comparative phylogenetic methods that can handle this new data are wanted. Here, we introduce the term \"phylogenetic latent variable models\" (PLVMs, pronounced \"plums\") for a class of models that has recently been used to infer the evolution of cellular states from systems-level molecular data, and develop a new parameterization and fitting strategy that is useful for comparative inference of biochemical networks. We deploy this new framework to infer the ancestral states and evolutionary dynamics of protein-interaction networks by analyzing >16,000 predominantly metazoan co-fractionation and affinity-purification mass spectrometry experiments. Based on these data, we estimate ancestral interactions across unikonts, broadly recovering protein complexes involved in translation, transcription, proteostasis, transport, and membrane trafficking. Using these results, we predict an ancient core of the Commander complex made up of CCDC22, CCDC93, C16orf62, and DSCR3, with more recent additions of COMMD-containing proteins in tetrapods. We also use simulations to develop model fitting strategies and discuss future model developments.

evolutionary biology

Network Identification Methods

Recently, network inference algorithms have grown tremendously in the field of systems biology because network identification is essential for understanding relationships between regulation mechanisms for genes, elucidating functional mechanisms underlying cellular processes, as well as identifying molecular targets for discoveries in medicines. This article provides a brief overview of different approaches used to identify biological networks and reviews recent advances in network identification.

Systems Biology

LM6-M: a high avidity rat monoclonal antibody to pectic α-1,5-L-arabinan

1,5-arabinan is an abundant structural feature of side chains of pectic rhamnogalacturonan-I which is a matrix constituent of plant cell walls. The study of arabinan in cells and tissues is driven by putative roles for this polysaccharide in the generation of cell wall and organ mechanical properties. The biological function(s) of arabinan is still uncertain and high quality molecular tools are required to detect its occurrence and monitor its dynamics. Here we report a new rat monoclonal antibody, LM6-M, similar in specificity to the published rat monoclonal antibody LM6 (Willats et al. (1998) Carbohydrate Research 308: 149-152). LM6-M is of the IgM immunoglobulin class and has a higher avidity for -1-5-L-arabinan than LM6. LM6-M displays high sensitivity in its detection of arabinan in in-vitro assays such as ELISA and epitope detection chromatography and in in-situ analyses.\n\nAbbreviations

plant biology

Multi-omic analysis of a hyper-diverse plant metabolic pathway reveals evolutionary routes to biological innovation

The diversity of life on Earth is a result of continual innovations in molecular networks influencing morphology and physiology. Plant specialized metabolism produces hundreds of thousands of compounds, offering striking examples of these innovations. To understand how this novelty is generated, we investigated the evolution of the Solanaceae family-specific, trichome-localized acylsugar biosynthetic pathway using a combination of mass spectrometry, RNA-seq, enzyme assays, RNAi and phylogenetics in non-model species. Our results reveal that hundreds of acylsugars are produced across the Solanaceae family and even within a single plant, revealing this phenotype to be hyper-diverse. The relatively short biosynthetic pathway experienced repeated cycles of innovation over the last 100 million years that include gene duplication and divergence, gene loss, evolution of substrate preference and promiscuity. This study provides mechanistic insights into the emergence of plant chemical novelty, and offers a template for investigating the [~]300,000 non-model plant species that remain underexplored.

genomics

A simple molecular mechanism explains multiple patterns of cell-size regulation

Increasingly accurate and massive data have recently shed light on the fundamental question of how cells maintain a stable size trajectory as they progress through the cell cycle. Microbes seem to use strategies ranging from a pure sizer, where the end of a given phase is triggered when the cell reaches a critical size, to pure adder, where the cell adds a constant size during a phase. Yet the biological origins of the observed spectrum of behavior remain elusive. We analyze a molecular size-control mechanism, based on experimental data from the yeast S. cerevisiae, that gives rise to behaviors smoothly interpolating between adder and sizer. The size-control is obtained from the titration of a repressor protein by an activator protein that accumulates more rapidly with increasing cell size. Strikingly, the size-control is composed of two different regimes: for small initial cell size, the size-control is a sizer, whereas for larger initial cell size, is is an imperfect adder. Our model thus indicates that the adder and critical size behaviors may just be different dynamical regimes of a single simple biophysical mechanism.

biophysics

Knowledge Formalization and High-Throughput Data Visualization Using Signaling Network Maps

Generation and usage of high-quality molecular signalling network maps can be augmented by standardising notations, establishing curation workflows and application of computational biology methods to exploit the knowledge contained in the maps. In this manuscript, we summarize the major aims and challenges of assembling information in the form of comprehensive maps of molecular interactions. Mainly, we share our experience gained while creating the Atlas of Cancer Signalling Network. In the step-by-step procedure, we describe the map construction process and suggest solutions for map complexity management by introducing a hierarchical modular map structure. In addition, we describe the NaviCell platform, a computational technology using Google Maps API to explore comprehensive molecular maps similar to geographical maps, and explain the advantages of semantic zooming principles for map navigation. We also provide the outline to prepare signalling network maps for navigation using the NaviCell platform. Finally, several examples of cancer high-throughput data analysis and visualization in the context of comprehensive signalling maps are presented.

systems biology

Predicting Causal Relationships from Biological Data: Applying Automated Casual Discovery on Mass Cytometry Data of Human Immune Cells

Learning the causal relationships that define a molecular system allows us to predict how the system will respond to different interventions. Distinguishing causality from mere association typically requires randomized experiments. Methods for automated causal discovery from limited experiments exist, but have so far rarely been tested in systems biology applications. In this work, we apply state-of-the art causal discovery methods on a large collection of public mass cytometry data sets, measuring intra-cellular signaling proteins of the human immune system and their response to several perturbations. We show how different experimental conditions can be used to facilitate causal discovery, and apply two fundamental methods that produce context-specific causal predictions. Causal predictions were reproducible across independent data sets from two different studies, but often disagree with the KEGG pathway databases. Within this context, we discuss the caveats we need to overcome for automated causal discovery to become a part of the routine data analysis in systems biology.

bioinformatics

The embryonic transcriptome of Arabidopsis thaliana

Cellular differentiation is associated with changes in transcript populations. Accurate quantification of transcriptomes during development can thus provide global insights into differentiation processes including the fundamental specification and differentiation events operating during plant embryogenesis. However, multiple technical challenges have limited the ability to obtain high quality early embryonic transcriptomes, namely the low amount of RNA obtainable and contamination from surrounding endosperm and seed-coat tissues. We compared the performance of three low-input mRNA sequencing (mRNA-seq) library preparation kits on 0.1 to 5 nanograms (ng) of total RNA isolated from Arabidopsis thaliana (Arabidopsis) embryos and identified a low-cost method with superior performance. This mRNA-seq method was then used to profile the transcriptomes of Arabidopsis embryos across eight developmental stages. By comprehensively comparing embryonic and post-embryonic transcriptomes, we found that embryonic transcriptomes do not resemble any other plant tissue we analyzed. Moreover, transcriptome clustering analyses revealed the presence of four distinct phases of embryogenesis which are enriched in specific biological processes. We also compared zygotic embryo transcriptomes with publicly available somatic embryo transcriptomes. Strikingly, we found little resemblance between zygotic embryos and somatic embryos derived from late-staged zygotic embryos suggesting that the molecular basis of somatic and zygotic embryogenesis are distinct from each other. In addition to the biological insights gained from our systematic characterization of the Arabidopsis embryonic transcriptome, we provide a data-rich resource for the community to explore.\n\nKey MessageArabidopsis embryos possess unique transcriptomes relative to other plant tissues including somatic embryos, and can be partitioned into four transcriptional phases with characteristic biological processes.

genomics

Embryonic Exposure to Valproic Acid Disrupts Social Predispositions in Newly-Hatched Chicks

Biological predispositions to attend to visual cues, such as those associated with face-like stimuli or with biological motion, guide social behavior from the first moments of life and have been documented in human neonates, infant monkeys and newly-hatched domestic chicks. In human neonates at high familial risk of Autism Spectrum Disorder (ASD), a lack of such predispositions has been recently reported. Prompted by these observations, we modeled ASD behavioral deficit in newborn chicks, using embryonic exposure to valproic acid (VPA), the histone deacetylases (HDACs) inhibitor that in humans is associated with an increased risk for developing ASD. We assessed spontaneous predispositions in newly-hatched, visually-naive chicks, by comparing responses to a stuffed hen vs. a scrambled version of it. We found that social predispositions were abolished in VPAtreated chicks. In contrast, experience-dependent learning mechanisms associated with filial imprinting were not affected. Our results indicate a specific effect of VPA on the development of biologically-predisposed social orienting mechanisms, opening new perspectives to investigate the molecular and neurobiological mechanisms involved in early ASD symptoms.

neuroscience

A simple and powerful analysis of lateral subdiffusion using single particle tracking

In biological membranes many factors such as cytoskeleton, lipid composition, crowding and molecular interactions deviate lateral diffusion from the expected random walks. These factors have different effects on diffusion but act simultaneously so the observed diffusion is a complex mixture of diffusive behaviors (directed, >Brownian, anomalous or confined). Therefore commonly used approaches to quantify diffusion based on averaging of the displacements, such as the mean square displacement, are not adapted to the analysis of this heterogeneity. We introduce a new parameter, the packing coefficient Pc, which gives an estimate of the degree of free movement that a molecule displays in a period of time independently of its global diffusivity. Applying this approach to two different situations (diffusion of a lipid probe and trapping of receptors at synapses), we show that Pc detected and localized temporary changes of diffusive behavior both in time and in space. More importantly, it allowed the detection of periods with very high confinement (~immobility), their frequency and duration, and thus it can be used to calculate the effective kon and koff of scaffolding interactions such those that immobilize receptors at synapses.

biophysics

Switch-like activation of Bruton’s tyrosine kinase by membrane-mediated dimerization

The transformation of molecular binding events into cellular decisions is the basis of most biological signal transduction. A fundamental challenge faced by these systems is that protein-ligand chemical affinities alone generally result in poor sensitivity to ligand concentration, endangering the system to error. Here, we examine the lipid-binding pleckstrin homology and Tec homology (PH-TH) module of Brutons tyrosine kinase (Btk) Using fluorescence correlation spectroscopy (FCS) and membrane-binding kinetic measurements, we identify a self-contained phosphatidylinositol (3,4,5)-trisphosphate (PIP3) sensing mechanism that achieves switch-like sensitivity to PIP3 levels, surpassing the intrinsic affinity discrimination of PIP3:PH binding. This mechanism employs multiple PIP3 binding as well as dimerization of Btk on the membrane surface. Mutational studies in live cells confirm that this mechanism is critical for activation of Btk in vivo. These results demonstrate how a single protein module can institute a minimalist coincidence detection mechanism to achieve high-precision discrimination of ligand concentration.

biophysics

The influence of X chromosome variants on trait neuroticism.

Autosomal variants have successfully been associated with trait neuroticism in genome-wide analysis of adequately-powered samples. But such studies have so far excluded the X chromosome from analysis. Here, we report genetic association analyses of X chromosome and XY pseudoautosomal single nucleotide polymorphisms (SNPs) and trait neuroticism using UK Biobank samples (N = 405,274). Significant association was found with neuroticism on the X chromosome for 204 markers found within three independent loci (a further 783 were suggestive). Most of these significant neuroticism-related X chromosome variants were located in intergenic regions (n = 713). Involvement of HS6ST2, which has been previously associated with sociability behaviour in the dog, was supported by single SNP and gene-based tests. We found that the amino acid and nucleotide sequences are highly conserved between dogs and humans. From the suggestive X chromosome variants, there were 19 nearby genes which could be linked to gene ontology information. Molecular function was primarily related to binding and catalytic activity; notable biological processes were cellular and metabolic, and nucleic acid binding and transcription factor protein classes were most commonly involved. X-variant heritability of neuroticism was estimated at 0.34% (SE = 0.07). A polygenic X-variant score created in an independent sample (maximum N {approx} 7300) did not predict significant variance in neuroticism, psychological distress, or depressive disorder. We conclude that the X chromosome harbours significant variants influencing neuroticism, and might prove important for other quantitative traits and complex disorders.

genetics

DNA methylation modules associate with incident cardiovascular disease and cumulative risk factor exposure

1 AbstractO_ST_ABSBackgroundC_ST_ABSEpigenome-wide association studies using DNA methylation have the potential to uncover novel biomarkers and mechanisms of cardiovascular disease (CVD) risk. However, the direction of causation for these associations is not always clear, and investigations to-date have generally failed to replicate at the level of individual loci. Here, we undertook module- and region-based DNA methylation analyses of incident CVD in the Womens Health Initiative (WHI) and Framingham Heart Study Offspring Cohort (FHS) in order to find more robust epigenetic biomarkers for cardiovascular risk.\n\nMethods and FindingsWe applied weighted gene correlation network analysis (WGCNA) and the Comb-p algorithm to find methylation modules and regions associated with incident CVD in the WHI dataset. We discovered two modules whose activation correlated with CVD risk and replicated across cohorts. One of these modules was enriched for development-related processes and overlaps strongly with epigenetic aging sites. For the other, we showed preliminary evidence for monocyte-specific effects and statistical links to cumulative exposure to traditional cardiovascular risk factors. Additionally, we found three regions (associated with the genes SLC9A1, SLC1A5, and TNRC6C) whose methylation associates with CVD risk.\n\nConclusionsIn sum, we present several epigenetic associations with incident CVD that reveal disease mechanisms related to development and monocyte biology. Furthermore, we show that epigenetic modules may act as a molecular readout of cumulative cardiovascular risk factor exposure, with implications for the improvement of clinical risk prediction.

genomics

A Statistical Procedure for Genome-wide Detection of QTL Hotspots Using Public Databases with Application to Rice

Genome-wide detection of quantitative trait loci (QTL) hotspots underlying variation in many molecular and phenotypic traits has been a key step in various biological studies since the QTL hotspots are highly informative and can be linked to the genes for the quantitative traits. Several statistical methods have been proposed to detect QTL hotspots. These hotspot detection methods rely heavily on permutation tests performed on summarized QTL data or individual-level data (with genotypes and phenotypes) from the genetical genomics experiments. In this article, we propose a statistical procedure for QTL hotspot detection by using the summarized QTL (interval) data collected in public web-accessible databases. First, a simple statistical method based on the uniform distribution is derived to convert the QTL interval data into the expected QTL frequency (EQF) matrix. And then, to account for the correlation structure among traits, the QTLs for correlated traits are grouped together into the same categories to form a reduced EQF matrix. Furthermore, a permutation algorithm on the EQF elements or on the QTL intervals is developed to compute a sliding scale of EQF thresholds, ranging from strict to liberal, for assessing the significance of QTL hotspots. With grouping, much stricter thresholds can be obtained to avoid the detection of spurious hotspots. Real example analysis and simulation study are carried out to illustrate our procedure, evaluate the performances and compare with other methods. It shows that our procedure can control the genome-wide error rates at the target levels, provide appropriate thresholds for correlated data and is comparable to the methods using individual-level data in hotspot detection. Depending on the thresholds used, more than 100 hotspots are detected in GRAMENE rice database. We also perform a genome-wide comparative analysis of the detected hotspots and the known genes collected in the Rice Q-TARO database. The comparative analysis reveals that the hotspots and genes are conformable in the sense that they co-localize closely and are functionally related to relevant traits. Our statistical procedure can provide a framework for exploring the networks among QTL hotspots, genes and quantitative traits in biological studies. The R codes that produce both numerical and graphical outputs of QTL hotspot detection in the genome are available on the worldwide web http://www.stat.sinica.edu.tw/~chkao/.

genetics