Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Molecular Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,225 records · Page 68Linked to original sources

High-Resolution Subtyping of Pediatric Low-Grade Glioma Using an Integrated Meta-Clustering Framework

Pediatric low-grade glioma (pLGG) is the most common type of brain tumor in children, accounting for approximately 30% of all central nervous system tumors in children. pLGG has multiple molecular subtypes that differ in disease progression, recurrence patterns, and treatment responses. Conventional wet lab approaches including molecular profiling and histopathological studies for pLGG characterization are time consuming, costly, and laborious. Recently, methods based on artificial intelligence (AI) or machine learning (ML) have been widely used for pLGG molecular categorization, but most of them can only identify two or three pLGG subtypes. To more comprehensively characterize the molecular subtypes of pLGG and their potential biological and therapeutic significance, we develop an integrated meta-clustering approach, namely Meta-pLGG, that can explore high resolution molecular subtypes and their transcriptional heterogeneity for pLGG. Specifically, we first performed multiple rounds of random projection (RP) to generate dimension-reduced feature vectors from pLGG transcriptomics data, each of which was subsequently clustered by different clustering algorithms including hierarchical clustering, K-means, Self-Organizing Maps (SOM), Non-negative Matrix Factorization (NMF), Gaussian Mixture Model (GMM), and Spectral Clustering, as base clustering methods. Then, to yield robust clustering performance, we integrated the clustering results of these RP based individual clustering algorithms by adopting a weighted meta-clustering (wMetaC) approach. Results based on 532 pLGG patients suggested that our proposed approach demonstrated superior stability and discriminative powers for higher resolution pLGG subtyping compared to conventional approaches. Based on consensus matrix analysis, we identified two major pLGG mega-subtypes, with one further subdivided into three subgroups and the other into two. Then, we performed cluster specific differential gene expression analysis, molecular pathway analysis, and gene-drug-disease association analysis. The results showed that the identified five subgroups exhibited significant subtype-specific transcriptomic heterogeneity. In summary, our meta-clustering approach demonstrated much higher performance and robustness in identifying higher resolution molecular subtypes of pLGG, revealing the molecular heterogeneity within pLGG and potentially providing new insights for more precise molecular subtyping and precision therapy.

bioinformatics

DIABLO - an integrative, multi-omics, multivariate method for multi-group classification

Systems biology approaches, leveraging multi-omics measurements, are needed to capture the complexity of biological networks while identifying the key molecular drivers of disease mechanisms. We present DIABLO, a novel integrative method to identify multi-omics biomarker panels that can discriminate between multiple phenotypic groups. In the multi-omics analyses of simulated and real-world datasets, DIABLO resulted in superior biological enrichment compared to other integrative methods, and achieved comparable predictive performance with existing multi-step classification schemes. DIABLO is a versatile approach that will benefit a diverse range of research areas, where multiple high dimensional datasets are available for the same set of specimens. DIABLO is implemented along with tools for model selection, and validation, as well as graphical outputs to assist in the interpretation of these integrative analyses (http://mixomics.org/).

Bioinformatics

RNA-binding activity of TRIM25 is mediated by its PRY/SPRY domain and is required for ubiquitination

TRIM25 is a novel RNA-binding protein and a member of the Tripartite Motif (TRIM) family of E3 ubiquitin ligases, which plays a pivotal role in the innate immune response. Almost nothing is known about its RNA-related roles in cell biology. Furthermore, its RNA-binding domain has not been characterized. Here, we reveal that RNA-binding activity of TRIM25 is mediated by its PRY/SPRY domain, which we postulate to be a novel RNA-binding domain. Using CLIP-seq and SILAC-based co-immunoprecipitation assays, we uncover TRIM25s endogenous RNA targets and protein binding partners. Finally, we show that the RNA-binding activity of TRIM25 is important for its ubiquitin ligase function. These results reveal new insights into the molecular roles and characteristics of RNA-binding E3 ubiquitin ligases and demonstrate that RNA could be an essential factor for their biological functions.

molecular biology

TESS: Bayesian inference of lineage diversification rates from (incompletely sampled) molecular phylogenies in R

SummaryMany fundamental questions in evolutionary biology entail estimating rates of lineage diversification (speciation - extinction). We develop a flexible Bayesian framework for specifying an effectively infinite array of diversification models--where rates are constant, vary continuously, or change episodically through time--and implement numerical methods to estimate parameters of these models from molecular phylogenies, even when species sampling is incomplete. Additionally we provide robust methods for comparing the relative and absolute fit of competing branching-process models to a given tree, thereby providing rigorous tests of biological hypotheses regarding patterns and processes of lineage diversification.\n\nAvailability and implementationthe source code for TESS is freely available at http://cran.r-project.org/web/packages/TESS/.\n\nContactSebastian.Hoehna@gmail.com

Bioinformatics

Bacteria: A novel source for potent mosquito feeding-deterrents

Antibiotic and insecticidal bioactivities of the extracellular secondary metabolites produced by entomopathogenic bacteria belonging to genus Xenorhabdus have been identified; however, their novel applications such as mosquito feeding-deterrence have not been reported. Here, we show that a mixture of compounds isolated from Xenorhabdus budapestensis in vitro cultures exhibits potent feeding-deterrent activity against three deadly mosquito vectors: Aedes aegypti, Anopheles gambiae and Culex pipiens. We further demonstrate that the deterrent-active fraction isolated from replicate bacterial cultures is consistently highly enriched in two modified peptides identical to the previously described fabclavines, strongly suggesting that these are molecular species responsible for feeding-deterrence. The mosquito feeding-deterrent activity in the fabclavines-rich fraction is comparable to or better than that of N, N-diethyl-3-methylbenzamide (also known as Deet) or picaridin in side-by-side assays. Our unique discovery lays the groundwork for research into biologically derived, peptide-based low molecular weight compounds isolated from bacteria for exploitation as mosquito repellents and feeding-deterrents.

microbiology

A two-step probing method to compare lysine accessibility across macromolecular complex conformations

Structural models of multi-megadalton molecular complexes are appearing in increasing numbers, in large part because of technical advances in cryo-electron microscopy realized over the last decade. However, the inherent complexity of large biological assemblies comprising dozens of components often limits the resolution of structural models. Furthermore, multiple functional configurations of a complex can leave a puzzle as to how one intermediate moves to the next stage. Orthogonal biochemical information is crucial to understanding the molecular interactions that drive those rearrangements. We present a two-step method for chemical probing detected by tandem mass-spectrometry to globally assess the reactivity of lysine residues within purified macromolecular complexes. Because lysine side chains often balance the negative charge of RNA in ribonucleoprotein complexes, the method is especially powerful for detecting changes in protein-RNA interactions. Probing the E. coli 30S ribosome subunit showed that the reactivity pattern of lysine residues quantitatively reflects structure models from X-ray crystallography. We assessed differences in two conformations of purified human spliceosomes. Our results demonstrate that this method supplies powerful biochemical information that aids in functional interpretation of atomic models of macromolecular complexes at the intermediate resolution often provided by cryo-electron microscopy.

molecular biology

Locoregional Radiogenomic Models Capture Gene Expression Heterogeneity in Glioblastoma

Radiogenomics mapping noninvasively determines important relationships between the molecular genotype and imaging phenotype of various tumors, allowing advances in both clinical care and cancer research. While early work has shown its technical feasibility, here we extend radiogenomic mapping to a locoregional level that can account for the molecular heterogeneity of tumors. To achieve this, our data processing pipeline relies on three main steps: 1) the use of multi-omics data fusion to generate a set of 100 interpretable gene modules, 2) the use of patch-based image analysis (specifically of contrast-enhanced T1-weighted weighted MR images) combined with Generalized Linear Models (GLM) to establish potential links between module expressions and local MR signal, and 3) the use of expression heatmaps based on GLMs decision values to explore visualization of tumor molecular heterogeneity. The performance of the proposed approach was evaluated using a leave-one-patient-out crossvalidation method as well as a separate validation data set. The top performing models were based on a small set of 20 features and yielded Area Under the receiver operating characteristic Curve (AUC) above 0.65 on the validation cohort for eight modules. Next, we demonstrate the clinical and biological interpretation of four modules using molecular expression heatmaps superimposed on clinical radiographic images, showing the potential for assessing tumor molecular heterogeneity and the utility of this method for precision treatment in clinical decision making and imaging surveillance.

bioinformatics

Logarithmic molecular sampling for next-generation sequencing

Next-generation sequencing enables measurement of chemical and biological signals at high throughput and falling cost. Conventional sequencing requires increasing sampling depth to improve signal to noise discrimination, a costly procedure that is also impossible when biological material is limiting. We introduce a new general sampling theory, Molecular Entropy encodinG (MEG), which uses biophysical principles to functionally encode molecular abundance before sampling. SeQUential DepletIon and enriCHment (SQUICH) is a specific example of MEG that, in theory and simulation, enables sampling at a logarithmic or better rate to achieve the same precision as attained with conventional sequencing. In proof-of-principle experiments, SQUICH reduces sequencing depth by a factor of 10. MEG is a general solution to a fundamental problem in molecular sampling and enables a new generation of efficient, precise molecular measurement at logarithmic or better sampling depth.

genomics

The E. coli molecular phenotype under different growth conditions

Modern systems biology requires extensive, carefully curated measurements of cellular components in response to different environmental conditions. While high-throughput methods have made transcriptomics and proteomics datasets widely accessible and relatively economical to generate, systematic measurements of both mRNA and protein abundances under a wide range of different conditions are still relatively rare. Here we present a detailed, genome-wide transcriptomics and proteomics dataset of E. coli grown under 34 different conditions. We manipulate concentrations of sodium and magnesium in the growth media, and we consider four different carbon sources glucose, gluconate, lactate, and glycerol. Moreover, samples are taken both in exponential and stationary phase, and we include two extensive time-courses, with multiple samples taken between 3 hours and 2 weeks. We find that exponential-phase samples systematically differ from stationary-phase samples, in particular at the level of mRNA. Regulatory responses to different carbon sources or salt stresses are more moderate, but we find numerous differentially expressed genes for growth on gluconate and under salt and magnesium stress. Our data set provides a rich resource for future computational modeling of E. coli gene regulation, transcription, and translation.

bioinformatics

Untargeted Mass Spectrometry-Based Metabolomics Tracks Molecular Changes in Raw and Processed Foods and Beverages

A major aspect of our daily lives is the need to acquire, store and prepare our food. Storage and preparation can have drastic effects on the compositional chemistry of our foods, but we have a limited understanding of the temporal nature of processes such as storage, spoilage, fermentation and brewing on the chemistry of the foods we eat. Here, we performed a temporal analysis of the chemical changes in foods during common household preparations using untargeted mass spectrometry and novel data analysis approaches. Common treatments of foods such as home fermentation of yogurt, brewing of tea, spoilage of meats and ripening of tomatoes altered the chemical makeup through time, through both chemical and biological processes. For example, brewing tea altered its composition by increasing the diversity of molecules, but this change was halted after 4 min of brewing. The results indicate that this is largely due to differential extraction of the material from the tea and not modification of the molecules during the brewing process. This is in contrast to the preparation of yogurt from milk, spoilage of meat and the ripening of tomatoes where biological transformations directly altered the foods molecular composition. Comprehensive assessment of chemical changes using multivariate statistics showed the varied impacts of the different food treatments, while analysis of individual chemical changes show specific alterations of chemical families in the different food types. The methods developed here represent novel approaches to studying the changes in food chemistry that can reveal global alterations in chemical profiles and specific transformations at the chemical level.\n\nO_LSTHighlightsC_LSTO_LIWe created a reference data set for tomato, milk to yogurt, tea, coffee, turkey and beef.\nC_LIO_LIWe show that normal preparation and handling affects the molecular make-up.\nC_LIO_LITea preparation is largely driven by differential extraction.\nC_LIO_LIFormation of yogurt involves chemical transformations.\nC_LIO_LIThe majority of meat molecules are not altered in 5 days at room temperature.\nC_LI

biochemistry

An integrative systems biology and experimental approach identifies convergence of epithelial plasticity, metabolism, and autophagy to promote chemoresistance

The evolution of therapeutic resistance is a major cause of death for patients with solid tumors. The development of therapy resistance is shaped by the ecological dynamics within the tumor microenvironment and the selective pressure induced by the host immune system. These ecological and selective forces often lead to evolutionary convergence on one or more pathways or hallmarks that drive progression. These hallmarks are, in turn, intimately linked to each other through gene expression networks. Thus, a deeper understanding of the evolutionary convergences that occur at the gene expression level could reveal vulnerabilities that could be targeted to treat therapy-resistant cancer. To this end, we used a combination of phylogenetic clustering, systems biology analyses, and wet-bench molecular experimentation to identify convergences in gene expression data onto common signaling pathways. We applied these methods to derive new insights about the networks at play during TGF-{beta}-mediated epithelial-mesenchymal transition in a lung cancer model system. Phylogenetics analyses of gene expression data from TGF-{beta} treated cells revealed evolutionary convergence of cells toward amine-metabolic pathways and autophagy during TGF-{beta} treatment. Using high-throughput drug screens, we found that knockdown of the autophagy regulatory, ATG16L1, re-sensitized lung cancer cells to cancer therapies following TGF-{beta}-induced resistance, implicating autophagy as a TGF-{beta}-mediated chemoresistance mechanism. Analysis of publicly-available clinical data sets validated the adverse prognostic importance of ATG16L expression in multiple cancer types including kidney, lung, and colon cancer patients. These analyses reveal the usefulness of combining evolutionary and systems biology methods with experimental validation to illuminate new therapeutic vulnerabilities.

cancer biology

Clinically Important sex differences in GBM biology revealed by analysis of male and female imaging, transcriptome and survival data

Sex differences in the incidence and outcome of human disease are broadly recognized but in most cases not adequately understood to enable sex-specific approaches to treatment. Glioblastoma (GBM), the most common malignant brain tumor, provides a case in point. Despite well-established differences in incidence, and emerging indications of differences in outcome, there are few insights that distinguish male and female GBM at the molecular level, or allow specific targeting of these biological differences. Here, using a quantitative imaging-based measure of response, we found that temozolomide chemotherapy is more effective in female compared to male GBM patients. We then applied a novel computational algorithm to linked GBM transcriptome and outcome data, and identified novel sex-specific molecular subtypes of GBM in which cell cycle and integrin signaling were identified as the critical determinants of survival for male and female patients, respectively. The clinical utility of cell cycle and integrin signaling pathway signatures was further established through correlations between gene expression and in vitro chemotherapy sensitivity in a panel of male and female patient-derived GBM cell lines. Together these results suggest that greater precision in GBM molecular subtyping can be achieved through sex-specific analyses, and that improved outcome for all patients might be accomplished via tailoring treatment to sex differences in molecular mechanisms.\n\nOne Sentence SummaryMale and female glioblastoma are biologically distinct and maximal chances for cure may require sex-specific approaches to treatment.

cancer biology

Sublethal effects of the neonicotinoid insecticide thiamethoxam on the transcriptome of the honeybee (Apis mellifera)

Neonicotinoid insecticides are now the most widely used insecticides in the world. Previous studies have indicated that sublethal doses of neonicotinoids impair learning, memory capacity, foraging and immunocompetence in honeybees (Apis mellifera). Despite this, few studies have been carried out on the molecular effects of neonicotinoids. In this study, we focus on the second-generation neonicotinoid thiamethoxam, which is currently widely used in agriculture to protect crops. Using high-throughput RNA-Seq, we investigated the transcriptome profile of honeybees after subchronic exposure to thiamethoxam (10 ppb) over 10 days. In total, 609 differentially-expressed genes (DEGs) were identified, of which 225 were up-regulated and 384 were down-regulated. The functions of some DEGs were identified, and GO enrichment analysis showed that the enriched DEGs were mainly linked to metabolism, biosynthesis and translation. KEGG pathway analysis showed that thiamethoxam affected biological processes including ribosomes, the oxidative phosphorylation pathway, tyrosine metabolism pathway, pentose and glucuronate interconversions and drug metabolism. Overall, our results provide a basis for understanding the molecular mechanisms of the complex interactions between neonicotinoid insecticides and honeybees.\n\nSummary statementNR1, Cyp6as5, nAChRa9 and nAChR{beta}2 were up-regulated in honeybees exposed to thiamethoxam, while CSP3, Obp21, defensin-1, Mrjp1, Mrjp3 and Mrjp4 were down-regulated.

molecular biology

Differential Community Detection in Paired Biological Networks

MotivationBiological networks unravel the inherent structure of molecular interactions which can lead to discovery of driver genes and meaningful pathways especially in cancer context. Often due to gene mutations, the gene expression undergoes changes and the corresponding gene regulatory network sustains some amount of localized re-wiring. The ability to identify significant changes in the interaction patterns caused by the progression of the disease can lead to the revelation of novel relevant signatures.\n\nMethodsThe task of identifying differential sub-networks in paired biological networks (A:control,B:case) can be re-phrased as one of finding dense communities in a single noisy differential topological (DT) graph constructed by taking absolute difference between the topological graphs of A and B. In this paper, we propose a fast two-stage approach, namely Differential Community Detection (DCD), to identify differential sub-networks as differential communities in a de-noised version of the DT graph. In the first stage, we iteratively re-order the nodes of the DT graph to determine approximate block diagonals present in the DT adjacency matrix using neighbourhood information of the nodes and Jaccard similarity. In the second stage, the ordered DT adjacency matrix is traversed along the diagonal to remove all the edges associated with a node, if that node has no immediate edges within a window. We then apply community detection methods on this de-noised DT graph to discover differential sub-networks as communities.\n\nResultsOur proposed DCD approach can effectively locate differential sub-networks in several simulated paired random-geometric networks and various paired scale-free graphs with different power-law exponents. The DCD approach easily outperforms community detection methods applied on the original noisy DT graph and recent statistical techniques in simulation studies. We applied DCD method on two real datasets: a) Ovarian cancer dataset to discover differential DNA co-methylation sub-networks in patients and controls; b) Glioma cancer dataset to discover the difference between the regulatory networks of IDH-mutant and IDH-wild-type. We demonstrate the potential benefits of DCD for finding network-inferred bio-markers/pathways associated with a trait of interest.\n\nConclusionThe proposed DCD approach overcomes the limitations of previous statistical techniques and the issues associated with identifying differential sub-networks by use of community detection methods on the noisy DT graph. This is reflected in the superior performance of the DCD method with respect to various metrics like Precision, Accuracy, Kappa and Specificity. The code implementing proposed DCD method is available at https://sites.google.com/site/ raghvendramallmlresearcher/codes.

systems biology

Transcriptomes and Raman spectra are linked linearly through a shared low-dimensional subspace

Raman spectroscopy is an imaging technique that can reflect whole-cell molecular compositions in vivo, and has been applied recently in cell biology to characterize different cell types and states. However, due to the complex molecular compositions and spectral overlaps, the interpretation of cellular Raman spectra have remained unclear. In this report, we compared cellular Raman spectra to transcriptomes of Schizosaccharomyces pombe and Escherichia coli, and provide firm evidence that they can be computationally connected and interpreted. Specifically, we find that the dimensions of high-dimensional Raman spectra and transcriptomes measured by RNA-seq can be effectively reduced and connected linearly through a shared low-dimensional subspace. Accordingly, we were able to reconstruct global gene expression profiles by applying the calculated transformation matrix to Raman spectra, and vice versa. Strikingly, highly expressed ncRNAs contributed to the Raman-transcriptome linear correspondence more significantly than mRNAs in S. pombe, which implies their major role in coordinating molecular compositions. This compatibility between whole-cell Raman spectra and transcriptomes marks an important and promising step towards establishing spectroscopic live-cell omics studies.

systems biology

On the Biological Signalling, Information and Estimation Limits of Birth Processes

Understanding and uncovering the mechanisms or motifs that molecular networks employ to regulate noise is a key problem in cell biology. As it is often difficult to obtain direct and detailed insight into these mechanisms, many studies instead focus on assessing the best precision attainable on the signalling pathways that compose these networks. Molecules signal one another over such pathways to solve noise regulating estimation and control problems. Quantifying the maximum precision of these solutions delimits what is achievable and allows hypotheses about underlying motifs to be tested without requiring detailed biological knowledge. The pathway capacity, which defines the maximum rate of transmitting information along it, is a widely used proxy for precision. Here it is shown, for estimation problems involving elementary yet biologically relevant birth-process networks, that capacity can be surprisingly misleading. A time-optimal signalling motif, called birth-following, is derived and proven to better the precision expected from the capacity, provided the maximum signalling rate constraint is large and the mean one above a certain threshold. When the maximum constraint is relaxed, perfect estimation is predicted by the capacity. However, the true achievable precision is found highly variable and sensitive to the mean constraint. Since the same capacity can map to different combinations of rate constraints, it can only equivocally measure precision. Deciphering the rate constraints on a signalling pathway may therefore be more important than computing its capacity.

systems biology

BacStalk: a comprehensive and interactive image analysis software tool for bacterial cell biology

Prokaryotes display a remarkable spatiotemporal organization of processes within individual cells. Investigations of the underlying mechanisms rely extensively on the analysis of microscopy images. Advanced image analysis software has revolutionized the cell-biological studies of established model organisms with largely symmetric rod-like cell shapes. However algorithms suitable for analyzing features of morphologically more complex model species are lacking although such unusually shaped organisms have emerged as treasure-troves of new molecular mechanisms and diversity in prokaryotic cell biology. To address this problem we developed BacStalk a simple interactive and easy-to-use MatLab-based software tool for quantitatively analyzing images of commonly and uncommonly shaped bacteria including stalked (budding) bacteria. BacStalk automatically detects the separate parts of the cells (cell body stalk bud or appendage) as well as their connections thereby allowing in-depth analyses of the organization of morphologically complex bacteria over time. BacStalk features the generation and visualization of concatenated fluorescence profiles along cells stalks appendages and buds to trace the spatiotemporal dynamics of fluorescent markers. Cells are interactively linked to demographs kymographs cell lineage analyses and scatterplots which enables intuitive and fast data exploration and thus significantly speeds up the image analysis process. Furthermore BacStalk introduces a 2D representation of demo- and kymographs enabling data representations in which the two spatial dimensions of the cell are preserved. The software was developed to handle large data sets and to generate publication-grade figures that can be easily edited. BacStalk therefore provides an advanced image analysis platform that extends the spectrum of model organisms for prokaryotic cell biology to bacteria with multiple morphologies and life cycles.\n\nIMPORTANCEProkaryotic cells show a striking degree of subcellular organization. Studies of the underlying mechanisms and their variation among different species greatly enhance our understanding of prokaryotic cell biology. The image analysis software tool BacStalk extracts an unprecedented amount of information from images of stalked bacteria, by generating interactive demographs, kymographs, cell lineages, and scatter plots that aid fast and thorough data analysis and representation. Notably, BacStalk can preserve the two spatial dimensions of cells when generating demographs and kymographs to accurately and intuitively reflect the intracellular organization. BacStalk also performs well on established, non-stalked model organisms with common or uncommon shapes. BacStalk therefore contributes to the advancement of prokaryotic cell biology, as it widens the spectrum of easily accessible model organisms and enables a more intuitive and interactive data analysis and visualization.

microbiology

A Bayesian network approach for modeling mixed features in TCGA ovarian cancer data

We propose an integrative framework to select important genetic and epigenetic features related to ovarian cancer and to quantify the causal relationships among these features using a logistic Bayesian network model based on The Cancer Genome Atlas data. The constructed Bayesian network has identified four gene clusters of distinct cellular functions, 13 driver genes, as well as some new biological pathways which may shed new light into the molecular mechanisms of ovarian cancer.

Systems Biology