Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Molecular Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,387 records · Page 77Linked to original sources

Sublethal effects of the neonicotinoid insecticide thiamethoxam on the transcriptome of the honeybee (Apis mellifera)

Neonicotinoid insecticides are now the most widely used insecticides in the world. Previous studies have indicated that sublethal doses of neonicotinoids impair learning, memory capacity, foraging and immunocompetence in honeybees (Apis mellifera). Despite this, few studies have been carried out on the molecular effects of neonicotinoids. In this study, we focus on the second-generation neonicotinoid thiamethoxam, which is currently widely used in agriculture to protect crops. Using high-throughput RNA-Seq, we investigated the transcriptome profile of honeybees after subchronic exposure to thiamethoxam (10 ppb) over 10 days. In total, 609 differentially-expressed genes (DEGs) were identified, of which 225 were up-regulated and 384 were down-regulated. The functions of some DEGs were identified, and GO enrichment analysis showed that the enriched DEGs were mainly linked to metabolism, biosynthesis and translation. KEGG pathway analysis showed that thiamethoxam affected biological processes including ribosomes, the oxidative phosphorylation pathway, tyrosine metabolism pathway, pentose and glucuronate interconversions and drug metabolism. Overall, our results provide a basis for understanding the molecular mechanisms of the complex interactions between neonicotinoid insecticides and honeybees.\n\nSummary statementNR1, Cyp6as5, nAChRa9 and nAChR{beta}2 were up-regulated in honeybees exposed to thiamethoxam, while CSP3, Obp21, defensin-1, Mrjp1, Mrjp3 and Mrjp4 were down-regulated.

molecular biology

Differential Community Detection in Paired Biological Networks

MotivationBiological networks unravel the inherent structure of molecular interactions which can lead to discovery of driver genes and meaningful pathways especially in cancer context. Often due to gene mutations, the gene expression undergoes changes and the corresponding gene regulatory network sustains some amount of localized re-wiring. The ability to identify significant changes in the interaction patterns caused by the progression of the disease can lead to the revelation of novel relevant signatures.\n\nMethodsThe task of identifying differential sub-networks in paired biological networks (A:control,B:case) can be re-phrased as one of finding dense communities in a single noisy differential topological (DT) graph constructed by taking absolute difference between the topological graphs of A and B. In this paper, we propose a fast two-stage approach, namely Differential Community Detection (DCD), to identify differential sub-networks as differential communities in a de-noised version of the DT graph. In the first stage, we iteratively re-order the nodes of the DT graph to determine approximate block diagonals present in the DT adjacency matrix using neighbourhood information of the nodes and Jaccard similarity. In the second stage, the ordered DT adjacency matrix is traversed along the diagonal to remove all the edges associated with a node, if that node has no immediate edges within a window. We then apply community detection methods on this de-noised DT graph to discover differential sub-networks as communities.\n\nResultsOur proposed DCD approach can effectively locate differential sub-networks in several simulated paired random-geometric networks and various paired scale-free graphs with different power-law exponents. The DCD approach easily outperforms community detection methods applied on the original noisy DT graph and recent statistical techniques in simulation studies. We applied DCD method on two real datasets: a) Ovarian cancer dataset to discover differential DNA co-methylation sub-networks in patients and controls; b) Glioma cancer dataset to discover the difference between the regulatory networks of IDH-mutant and IDH-wild-type. We demonstrate the potential benefits of DCD for finding network-inferred bio-markers/pathways associated with a trait of interest.\n\nConclusionThe proposed DCD approach overcomes the limitations of previous statistical techniques and the issues associated with identifying differential sub-networks by use of community detection methods on the noisy DT graph. This is reflected in the superior performance of the DCD method with respect to various metrics like Precision, Accuracy, Kappa and Specificity. The code implementing proposed DCD method is available at https://sites.google.com/site/ raghvendramallmlresearcher/codes.

systems biology

Transcriptomes and Raman spectra are linked linearly through a shared low-dimensional subspace

Raman spectroscopy is an imaging technique that can reflect whole-cell molecular compositions in vivo, and has been applied recently in cell biology to characterize different cell types and states. However, due to the complex molecular compositions and spectral overlaps, the interpretation of cellular Raman spectra have remained unclear. In this report, we compared cellular Raman spectra to transcriptomes of Schizosaccharomyces pombe and Escherichia coli, and provide firm evidence that they can be computationally connected and interpreted. Specifically, we find that the dimensions of high-dimensional Raman spectra and transcriptomes measured by RNA-seq can be effectively reduced and connected linearly through a shared low-dimensional subspace. Accordingly, we were able to reconstruct global gene expression profiles by applying the calculated transformation matrix to Raman spectra, and vice versa. Strikingly, highly expressed ncRNAs contributed to the Raman-transcriptome linear correspondence more significantly than mRNAs in S. pombe, which implies their major role in coordinating molecular compositions. This compatibility between whole-cell Raman spectra and transcriptomes marks an important and promising step towards establishing spectroscopic live-cell omics studies.

systems biology

On the Biological Signalling, Information and Estimation Limits of Birth Processes

Understanding and uncovering the mechanisms or motifs that molecular networks employ to regulate noise is a key problem in cell biology. As it is often difficult to obtain direct and detailed insight into these mechanisms, many studies instead focus on assessing the best precision attainable on the signalling pathways that compose these networks. Molecules signal one another over such pathways to solve noise regulating estimation and control problems. Quantifying the maximum precision of these solutions delimits what is achievable and allows hypotheses about underlying motifs to be tested without requiring detailed biological knowledge. The pathway capacity, which defines the maximum rate of transmitting information along it, is a widely used proxy for precision. Here it is shown, for estimation problems involving elementary yet biologically relevant birth-process networks, that capacity can be surprisingly misleading. A time-optimal signalling motif, called birth-following, is derived and proven to better the precision expected from the capacity, provided the maximum signalling rate constraint is large and the mean one above a certain threshold. When the maximum constraint is relaxed, perfect estimation is predicted by the capacity. However, the true achievable precision is found highly variable and sensitive to the mean constraint. Since the same capacity can map to different combinations of rate constraints, it can only equivocally measure precision. Deciphering the rate constraints on a signalling pathway may therefore be more important than computing its capacity.

systems biology

BacStalk: a comprehensive and interactive image analysis software tool for bacterial cell biology

Prokaryotes display a remarkable spatiotemporal organization of processes within individual cells. Investigations of the underlying mechanisms rely extensively on the analysis of microscopy images. Advanced image analysis software has revolutionized the cell-biological studies of established model organisms with largely symmetric rod-like cell shapes. However algorithms suitable for analyzing features of morphologically more complex model species are lacking although such unusually shaped organisms have emerged as treasure-troves of new molecular mechanisms and diversity in prokaryotic cell biology. To address this problem we developed BacStalk a simple interactive and easy-to-use MatLab-based software tool for quantitatively analyzing images of commonly and uncommonly shaped bacteria including stalked (budding) bacteria. BacStalk automatically detects the separate parts of the cells (cell body stalk bud or appendage) as well as their connections thereby allowing in-depth analyses of the organization of morphologically complex bacteria over time. BacStalk features the generation and visualization of concatenated fluorescence profiles along cells stalks appendages and buds to trace the spatiotemporal dynamics of fluorescent markers. Cells are interactively linked to demographs kymographs cell lineage analyses and scatterplots which enables intuitive and fast data exploration and thus significantly speeds up the image analysis process. Furthermore BacStalk introduces a 2D representation of demo- and kymographs enabling data representations in which the two spatial dimensions of the cell are preserved. The software was developed to handle large data sets and to generate publication-grade figures that can be easily edited. BacStalk therefore provides an advanced image analysis platform that extends the spectrum of model organisms for prokaryotic cell biology to bacteria with multiple morphologies and life cycles.\n\nIMPORTANCEProkaryotic cells show a striking degree of subcellular organization. Studies of the underlying mechanisms and their variation among different species greatly enhance our understanding of prokaryotic cell biology. The image analysis software tool BacStalk extracts an unprecedented amount of information from images of stalked bacteria, by generating interactive demographs, kymographs, cell lineages, and scatter plots that aid fast and thorough data analysis and representation. Notably, BacStalk can preserve the two spatial dimensions of cells when generating demographs and kymographs to accurately and intuitively reflect the intracellular organization. BacStalk also performs well on established, non-stalked model organisms with common or uncommon shapes. BacStalk therefore contributes to the advancement of prokaryotic cell biology, as it widens the spectrum of easily accessible model organisms and enables a more intuitive and interactive data analysis and visualization.

microbiology

A Bayesian network approach for modeling mixed features in TCGA ovarian cancer data

We propose an integrative framework to select important genetic and epigenetic features related to ovarian cancer and to quantify the causal relationships among these features using a logistic Bayesian network model based on The Cancer Genome Atlas data. The constructed Bayesian network has identified four gene clusters of distinct cellular functions, 13 driver genes, as well as some new biological pathways which may shed new light into the molecular mechanisms of ovarian cancer.

Systems Biology

Systematic analysis of mouse genome reveals distinct evolutionary and functional properties among circadian and ultradian genes

In living organisms, biological clocks regulate 24 h (circadian) molecular, physiological, and behavioral rhythms to maintain homeostasis and synchrony with predictable environmental changes. Harmonics of these circadian rhythms having periods of 8 hours and 12 hours (ultradian) have been documented in several species. In mouse liver, harmonics of the 24-hour period of gene transcription hallmarked genes oscillating with a frequency two or three times faster than the circadian circuitry. Many of these harmonic transcripts enriched pathways regulating responses to environmental stress and coinciding preferentially with subjective dawn and dusk. We hypothesized that these stress anticipatory genes would be more evolutionarily conserved than background circadian and non-circadian genes. To investigate this issue, we performed broad computational analyses of genes/proteins oscillating at different frequency ranges across several species and showed that ultradian genes/proteins, especially those oscillating with a 12-hour periodicity, are more likely to be of ancient origin and essential in mice. In summary, our results show that genes with ultradian transcriptional patterns are more likely to be phylogenetically conserved and associated with the primeval and inevitable dawn/dusk transitions.

bioinformatics

Spatio-Temporal Network Dynamics of Genes Underlying Schizophrenia

Schizophrenia (SZ) is a debilitating mental illness with multigenic etiology and high heritability. Despite extensive genetic studies the molecular etiology stays enigmatic. A systems biology study had suggested a protein-protein interaction (PPI) network for SZ with 504 novel PPIs amongst which several genes happen to be drug targets of existing FDA approved drugs. Although the PPI network presented all possible pairs of interactions (known and novel), it lacks a spatio-temporal information. The onset of psychiatric disorders is predominantly in adolescent and young adult stages, often accompanied by subtle structural abnormalities in multiple regions of the brain. Hence, there is a need to redefine the generic PPI network as a function of time (developmental stages) and space (brain regions). The availability of BrainSpan atlas data allowed us to redefine the SZ interactome as a function of space and time. The absence of non-synonymous variants in centenarians and non-psychiatric ExAC database allowed us to identify the variants of criticality. The expression of candidate genes in different brain regions and during developmental stages, responsible for cognitive processes as well as the onset of disease were studied. A subset of novel interactors detected in the network was further validated using gene-expression data of psychiatric postmortem brains. From the long list of drug targets proposed from the interactome study and based on the microarray gene-expression results, we have shortlisted a probable subset of 10 drug targets (targeted by 34 FDA approved drugs) coalescing into 81 biological pathways, that could be potentially repurposed for neuropsychiatric disorders.

bioinformatics

in vitro egg production by the human parasite Schistosoma mansoni

Schistosomes infect over 200 million people. The prodigious egg output of these parasites is the sole driver of pathology due to infection, yet our understanding of their sexual reproduction is limited because egg production is not sustained for more than a few days in vitro. Here, we describe culture conditions that support schistosome sexual development and sustained egg production in vitro. Female schistosomes rely on continuous pairing with male worms to fuel the maturation of their reproductive organs. Exploiting these new culture conditions, we explore the process of male-stimulated female maturation and demonstrate that physical contact with a male worm, and not insemination, is sufficient to induce female development and the production of viable parthenogenetic haploid embryos. We further report the characterization of a novel nuclear receptor, that we call vitellogenic factor 1, that is essential for female sexual development following pairing with a male worm. Taken together, these results provide a platform to study the fascinating sexual biology of these parasites on a molecular level, illuminating new strategies to control schistosome egg production.

microbiology

Improved genome assembly and annotation for the rock pigeon (Columba livia)

The domestic rock pigeon (Columba livia) is among the most widely distributed and phenotypically diverse avian species. This species is broadly studied in ecology, genetics, physiology, behavior, and evolutionary biology, and has recently emerged as a model for understanding the molecular basis of anatomical diversity, the magnetic sense, and other key aspects of avian biology. Here we report an update to the C. livia genome reference assembly and gene annotation dataset. Greatly increased scaffold lengths in the updated reference assembly, along with an updated annotation set, provide improved tools for evolutionary and functional genetic studies of the pigeon, and for comparative avian genomics in general.

genomics

Reverse enGENEering of regulatory networks from Big Data: a guide for a biologist

Omics technologies enable unbiased investigation of biological systems through massively parallel sequence acquisition or molecular measurements, bringing the life sciences into the era of Big Data. A central challenge posed by such omics datasets is how to transform this data into biological knowledge. For example, how to use this data to answer questions such as: which functional pathways are involved in cell differentiation? Which genes should we target to stop cancer? Network analysis is a powerful and general approach to solve this problem consisting of two fundamental stages, network reconstruction and network interrogation. Herein, we provide an overview of network analysis including a step by step guide on how to perform and use this approach to investigate a biological question. In this guide, we also include the software packages that we and others employ for each of the steps of a network analysis workflow.

Systems Biology

A comprehensive, mechanistically detailed, and executable model of the Cell Division Cycle in Saccharomyces cerevisiae

Understanding how cellular functions emerge from the underlying molecular mechanisms is a key challenge in biology. This will require computational models, whose predictive power is expected to increase with coverage and precision of formulation. Genome-scale models revolutionised the metabolic field and made the first whole-cell model possible. However, the lack of genome-scale models of signalling networks blocks the development of eukaryotic whole-cell models. Here, we present a comprehensive mechanistic model of the molecular network that controls the cell division cycle in Saccharomyces cerevisiae. We use rxncon, the reaction-contingency language, to neutralise the scalability issues preventing formulation, visualisation and simulation of signalling networks at the genome-scale. We use parameter-free modelling to validate the network and to predict genotype-to-phenotype relationships down to residue resolution. This mechanistic genome-scale model offers a new perspective on eukaryotic cell cycle control, and opens up for similar models - and eventually whole-cell models - of human cells.

systems biology

Species limits in butterflies (Lepidoptera: Nymphalidae): Reconciling classical taxonomy with the multispecies coalescent

Species delimitation is at the core of biological sciences. During the last decade, molecular-based approaches have advanced the field by providing additional sources of evidence to classical, morphology-based taxonomy. However, taxonomy has not yet fully embraced molecular species delimitation beyond threshold-based, single-gene approaches, and taxonomic knowledge is not commonly integrated to multi-locus species delimitation models. Here we aim to bridge empirical data (taxonomic and genetic) with recently developed coalescent-based species delimitation approaches. We use the multispecies coalescent model as implemented in two Bayesian methods (DISSECT/STACEY and BP&P) to infer species hypotheses. In both cases, we account for phylogenetic uncertainty (by not using any guide tree) and taxonomic uncertainty (by measuring the impact of using or not a priori taxonomic assignment to specimens). We focus on an entire Neotropical tribe of butterflies, the Haeterini (Nymphalidae: Satyrinae). We contrast divergent taxonomic opinion--splitting, lumping and misclassifying species--in the light of different phenotypic classifications proposed to date. Our results provide a solid background for the recognition of 22 species. The synergistic approach presented here overcomes limitations in both traditional taxonomy (e.g. by recognizing cryptic species) and molecular-based methods (e.g. by recognizing structured populations, and not raise them to species). Our framework provides a step forward towards standardization and increasing reproducibility of species delimitations.

evolutionary biology

Sirt7 regulates circadian phase coherence of hepatic circadian clock via a body temperature/Hsp70-Sirt7-Cry1 axis

The biological clock is generated in the hypothalamic suprachiasmatic nucleus (SCN), which synchronizes peripheral oscillators to coordinate physiological and behavioral activities throughout the body. Disturbance of circadian phase coherence between the central and peripheral could disrupt rhythms and thus cause diseases and aging. Here, we identified hepatic Sirt7 as an early element responsive to light, which ensures the phase coherence in mouse liver. Loss of Sirt7 leads to advanced liver circadian phase; restricted feeding in daytime entrains hepatic clock more rapidly in Sirt7-/- mice compared to wild-types. Molecularly, a light-driven body temperature (BT) oscillation induces rhythmic expression of Hsp70, which binds to and promotes the ubiquitination and proteasomal degradation of Sirt7. Sirt7 rhythmically deacetylates Cry1 on K565/579 and promotes Fbxl3-mediated degradation, thus coupling hepatic clock to the central pacemaker. Together, our data identify a novel BT/Hsp70-Sirt7-Cry1 axis, which transmits biological timing cues from the central to the peripheral and ensures circadian phase coherence in livers.

molecular biology

Network-based prediction of protein interactions

As biological function emerges through interactions between a cells molecular constituents, understanding cellular mechanisms requires us to catalogue all physical interactions between proteins [1-4]. Despite spectacular advances in high-throughput mapping, the number of missing human protein-protein interactions (PPIs) continues to exceed the experimentally documented interactions [5, 6]. Computational tools that exploit structural, sequence or network topology information are increasingly used to fill in the gap, using the patterns of the already known interactome to predict undetected, yet biologically relevant interactions [7-9]. Such network-based link prediction tools rely on the Triadic Closure Principle (TCP) [10-12], stating that two proteins likely interact if they share multiple interaction partners. TCP is rooted in social network analysis, namely the observation that the more common friends two individuals have, the more likely that they know each other [13, 14]. Here, we offer direct empirical evidence across multiple datasets and organisms that, despite its dominant use in biological link prediction, TCP is not valid for most protein pairs. We show that this failure is fundamental - TCP violates both structural constraints and evolutionary processes. This understanding allows us to propose a link prediction principle, consistent with both structural and evo-lutionary arguments, that predicts yet uncovered protein interactions based on paths of length three (L3). A systematic computational cross-validation shows that the L3 principle significantly outperforms existing link prediction methods. To experimentally test the L3 predictions, we perform both large-scale high-throughput and pairwise tests, finding that the predicted links test positively at the same rate as previously known interactions, suggesting that most (if not all) predicted interactions are real. Combining L3 predictions with experimen-tal tests provided new interaction partners of FAM161A, a protein linked to retinitis pigmentosa, offering novel insights into the molecular mechanisms that lead to the disease. Because L3 is rooted in a fundamental biological principle, we expect it to have a broad applicability, enabling us to better understand the emergence of biological function under both healthy and pathological conditions.\n\nSummaryWe unveil a fundamental organizing principle of biological networks and demonstrate its predictive power for uncovering novel protein interactions.

systems biology

iterative Random Forests to discover predictive and stable high order interactions

Genomics has revolutionized biology, enabling the interrogation of whole transcriptomes, genome-wide binding sites for proteins, and many other molecular processes. However, individual genomic assays measure elements that interact in vivo as components of larger molecular machines. Understanding how these high-order interactions drive gene expression presents a substantial statistical challenge. Building on Random Forests (RF), Random Intersection Trees (RITs), and through extensive, biologically inspired simulations, we developed the iterative Random Forest algorithm (iRF). iRF trains a feature-weighted ensemble of decision trees to detect stable, high-order interactions with same order of computational cost as RF. We demonstrate the utility of iRF for high-order interaction discovery in two prediction problems: enhancer activity in the early Drosophila embryo and alternative splicing of primary transcripts in human derived cell lines. In Drosophila, among the 20 pairwise transcription factor interactions iRF identifies as stable (returned in more than half of bootstrap replicates), 80% have been previously reported as physical interactions. Moreover, novel third-order interactions, e.g. between Zelda (Zld), Giant (Gt), and Twist (Twi), suggest high-order relationships that are candidates for follow-up experiments. In human-derived cells, iRF re-discovered a central role of H3K36me3 in chromatin-mediated splicing regulation, and identified novel 5th and 6th order interactions, indicative of multi-valent nucleosomes with specific roles in splicing regulation. By decoupling the order of interactions from the computational cost of identification, iRF opens new avenues of inquiry into the molecular mechanisms underlying genome biology.

genomics

Receptor crosstalk improves concentration sensing of multiple ligands

Cells need to reliably sense external ligand concentrations to achieve various biological functions such as chemotaxis or signaling. The molecular recognition of ligands by surface receptors is degenerate in many systems leading to crosstalk between different receptors. Crosstalk is often thought of as a deviation from optimal specific recognition, as the binding of non-cognate ligands can interfere with the detection of the receptors cognate ligand, possibly leading to a false triggering of a downstream signaling pathway. Here we quantify the optimal precision of sensing the concentrations of multiple ligands by a collection of promiscuous receptors. We demonstrate that crosstalk can improve precision in concentration sensing and discrimination tasks. To achieve superior precision, the additional information about ligand concentrations contained in short binding events of the noncognate ligand should be exploited. We present a proofreading scheme to realize an approximate estimation of multiple ligand concentrations that reaches a precision close to the derived optimal bounds. Our results help rationalize the observed ubiquity of receptor crosstalk in molecular sensing.

cell biology

A Hypothesis to Explain Cancers in Confined Colonies of Naked Mole Rats

Naked mole rats (NMRs) are subterranean eusocial mammals, known for their virtual absence of aging in their first 20 to 30 years of life, and their apparent resistance to cancer development. As such, this species has become an important biological model for investigating the physiological and molecular mechanisms behind cancer resistance. Two recent studies have discovered middle and late-aged worker (that is, non-breeding) NMRs in captive populations exhibiting neoplasms, consistent with cancer development, challenging the claim that NMRs are cancer resistant. These cases are possibly artefacts of inbreeding or certain rearing conditions in captivity, but they are also consistent with evolutionary theory.\n\nWe present field data showing that worker NMRs live on average for 1 to 2 years. This, together with considerable knowledge about the biology of this species, provides the basis for an evolutionary explanation for why debilitating cancers in NMRs should be rare in captive populations and absent in the wild. Whereas workers are important for maintaining tunnels, colony defence, brood care, and foraging, they are highly vulnerable to predation. However, surviving workers either replace dead breeders, or assume other less active functions whilst preparing for possible dispersal. These countervailing forces (selection resulting in aging due to early-life investments in worker function, and selection for breeder longevity) along with the fact that all breeders derive from the worker morph, can explain the low levels of cancer observed by these recent studies in captive colonies. Because workers in the field typically never reach ages where cancer becomes a risk to performance or mortality, those rare observations of neoplastic growth should be confined to the artificial environments where workers survive to ages rarely if ever occurring in the wild. Thus, we predict that the worker phenotype fortuitously benefits from anti-aging and cancer protection in captive populations.

Cancer Biology