Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Molecular Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,135 records · Page 63Linked to original sources

Targeted RNA NextGenSeq profiling in oncology using single molecule molecular inversion probes.

Hundreds of biology-based precision drugs are available that neutralize aberrant molecular pathways in cancer. Molecular heterogeneity and the lack of reliable companion diagnostic biomarkers for many drugs makes targeted treatment of cancer inaccurate for many individuals, leading to futile overtreatment. To acquire a comprehensive insight in aberrant actionable biological pathways in individual cancers we applied a cost-effective targeted RNA next generation sequencing (NGS) technique. The test allows NGS-based measurement of transcript levels and splice variants of hundreds of genes with established roles in the biological behavior in many cancer types. We here present proof of concept that the technique generates a correct molecular diagnosis and a prognosis for glioma patients. The test not only confirmed known brain cancer-associated molecular aberrations but also identified aberrant expression levels of actionable genes and mutations that are associated with other cancer types. Targeted RNA-NGS is therefore a highly attractive method to guide precision therapy for the individual patient based on pathway analysis.

cancer biology

Tracking stem cell differentiation without biomarkers using pattern recognition and phase contrast imaging

Bio-image informatics is the systematic application of image analysis algorithms to large image datasets to provide an objective method for accurately and consistently scoring image data. Within this field, pattern recognition (PR) is a form of supervised machine learning where the computer identifies relevant patterns in groups (classes) of images after being trained on examples. Rather than segmentation, image-specific algorithms or adjustable parameter sets, PR relies on extracting a common set of image descriptors (features) from the entire image to determine similarities and differences between image classes.\n\nGross morphology can be the only available description of biological systems prior to their molecular characterization, but these descriptions can be subjective and qualitative. In principle, generalized PR can provide an objective and quantitative characterization of gross morphology, thus providing a means of computationally defining morphological biomarkers. In this study, we investigated the potential of a pattern recognition approach to a problem traditionally addressed using genetic or biochemical biomarkers. Often these molecular biomarkers are unavailable for investigating biological processes that are not well characterized, such as the initial steps of stem cell differentiation.\n\nHere we use a general contrast technique combined with generalized PR software to detect subtle differences in cellular morphology present in early differentiation events in murine embryonic stem cells (mESC) induced to differentiate by the overexpression of selected transcription factors. Without the use of reporters, or a priori knowledge of the relevant morphological characteristics, we identified the earliest differentiation event (3 days), reproducibly distinguished eight morphological trajectories, and correlated morphological trajectories of 40 mESC clones with previous micro-array data. Interestingly, the six transcription factors that caused the greatest morphological divergence from an ESC-like state were previously shown by expression profiling to have the greatest influence on the expression of downstream genes.

cell biology

A Method To Assess Significance Of Differences In RNA Expression Levels Among Specific Groups Of Genes

Genome-wide molecular gene expression studies generally compare expression values for each gene across multiple conditions followed by cluster and gene set enrichment analysis to determine whether differentially expressed genes are enriched in specific biochemical pathways, cellular components, biological processes, and/or molecular functions, etc. This approach to analyzing differences in gene expression enables discovery of gene function, but is not useful to determine whether pre-defined groups of genes share or diverge in their expression patterns in response to treatments nor to assess the correctness of pre-defined gene set groupings. Here we present a simple method that changes the dimension of comparison by treating genes as variable traits to directly assess significance of differences in expression levels among pre-defined gene groups. Because expression distributions are typically skewed (thus unfit for direct assessment using Gaussian statistical methods) our method involves transforming expression data to approximate a normal distribution followed by dividing the genes into groups, then applying Gaussian parametric methods to assess significance of observed differences. This method enables the assessment of differences in gene expression distributions within and across samples, enabling hypothesis-based comparison among groups of genes. We demonstrate this method by assessing the significance of specific gene groups differential response to heat stress conditions in maize.\n\nAbbreviations

bioinformatics

Centrality and the shortest path approach on the human interactome

A network is one of the most convenient way to represent interactions between biological entities in systems biology. A network of molecular interactions is a graph in which the vertices are biological component and the edges correspond to the interactions between them. Many notions and approaches for network analysis came to systems biology from the theory of graphs - a field of mathematics that study graphs. We focused on the study of the shortest path approach in this work. We investigated whether this approach yields valid molecular paths. To perform this, the shortest paths in the human interactome (derived from HPRD and HIPPIE databases) were found between all relevant combinations of proteins taken from eight well-studied highly conserved signaling pathways from the KEGG database (NF-kappa B, MAPK, Jak-STAT, mTOR, ErbB, Wnt, TGF-beta and the signaling part of the apoptotic process). Canonical paths were systematically compared with the shortest counterparts and centrality of vertices and paths in subnetworks induced by the shortest paths were analyzed. We found that the sets of the shortest paths contain the canonical counterparts only for very short canonical paths (length 2-3 interactions). We also found that high centrality vertices tend to belong to canonical counterparts, to less extent this can also be said about high centrality paths.

systems biology

Exploring genes of rectal cancer for new treatments based on protein interaction network

ObjectiveTo develop a protein-protein interaction network of rectal cancer, which is based on genetic genes as well as to predict biological pathways underlying the molecular complexes in the network. In order to analyze and summarize genetic markers related to diagnosis and prognosis of rectal cancer.\n\nMethodsthe genes expression profile was downloaded from OMIM (Online Mendelian Inheritance in Man) database; the protein-protein interaction network of rectal cancer was established by Cytoscape; the molecular complexes in the network were detected by Clusterviz plugin and the pathways enrichment of molecular complexes were performed by DAVID online and Bingo (The Biological Networks Gene Ontology tool).\n\nResults and DiscussionA total of 127 rectal cancer genes were identified to differentially express in OMIM Database. The protein-protein interaction network of rectal cancer was contained 966 nodes (proteins), 3377 edges (interactive relationships) and 7 molecular complexes (score>7.0). Regulatory effects of genes and proteins were focused on cell cycle, transcription regulation and cellular protein metabolic process. Genes of DDK1, sparcl1, wisp2, cux1, pabpc1, ptk2 and htral were significant nodes in PPI network. The discovery of featured genes which were probably related to rectal cancer, has a great significance on studying mechanism, distinguishing normal and cancer tissues, and exploring new treatments for rectal cancer.

Cancer Biology

Biological Organization, Biological Information, and Knowledge

A concept of information designed to handle information conveyed by organizations is introduced. This concept of information may be used at all biological scales: from molecular and intracellular to multi-cellular organisms and human beings, and further on to collectivities, societies and culture. In this short account, two ground concepts necessary for developing the definition will also be introduced: whole-part graphs, a model for biological organization, and synexions, their immersion into space-time. This definition of information formalizes perception, observers and interpretation; allowing for considering information-exchange as a basic form of biological interaction. Some of its elements will be clarified by arguing and explaining why the immersion of whole-part graphs in (the physical) space-time is needed.

Systems Biology

Ancestral Reconstruction of Gene Blocks using an Event-Based Method

Complexity is a fundamental attribute of life. Complex systems are made of parts that together perform functions that a single component, or subsets containing individual components, cannot. Examples of complex molecular systems include protein structures such as the F1Fo-ATPase, the ribosome, or the flagellar motor: each one of these structures requires most or all of its components to function properly. Given the ubiquity of complex systems in the biosphere, understanding the evolution of complexity is central to biology. At the molecular level, operons are a classic example of a complex system. An operons genes are co-transcribed under the control of a single promoter to a polycistronic mRNA molecule, and the operons gene products often form molecular complexes or metabolic pathways. With the large number of complete bacterial genomes available, we now have the opportunity to explore the evolution of these complex entities, by identifying possible intermediate states of operons. In this work, we developed a maximum parsimony algorithm to reconstruct ancestral operon states, and show a simple vertical evolution model of how operons may evolve from the individual component genes. We describe several ancestral states that are plausible functional intermediate forms leading to the full operon. We also offer Reconstruction of Ancestral Gene blocks Using Events or ROAGUE as a software tool for those interested in exploring gene block and operon evolution.\n\nThe software accompanying this paper is available under GPLv3 license in: https://github.com/nguyenngochuy91/Ancestral-Blocks-Reconstruction.\n\nAll figures in this paper are available in enlarged downloadable form from:https://github.com/nguyenngochuy91/Ancestral-Blocks-Reconstruction/tree/master/images

evolutionary biology

Multi-Omics factor analysis disentangles heterogeneity in blood cancer

Multi-omic studies promise the improved characterization of biological processes across molecular layers. However, methods for the unsupervised integration of the resulting heterogeneous datasets are lacking. We present Multi-Omics Factor Analysis (MOFA), a computational method for discovering the principal sources of variation in multi-omic datasets. MOFA infers a set of (hidden) factors that capture biological and technical sources of variability. It disentangles axes of heterogeneity that are shared across multiple modalities and those specific to individual data modalities. The learnt factors enable a variety of downstream analyses, including identification of sample subgroups, data imputation, and the detection of outlier samples. We applied MOFA to a cohort of 200 patient samples of chronic lymphocytic leukaemia, profiled for somatic mutations, RNA expression, DNA methylation and ex-vivo drug responses. MOFA identified major dimensions of disease heterogeneity, including immunoglobulin heavy chain variable region status, trisomy of chromosome 12 and previously underappreciated drivers, such as response to oxidative stress. In a second application, we used MOFA to analyse single-cell multiomics data, identifying coordinated transcriptional and epigenetic changes along cell differentiation.

bioinformatics

Modelling genotypes in their physical microenvironment to predict single- and multi-cellular behaviour

A cells phenotype is the set of observable characteristics resulting from the interaction of the genotype with the surrounding environment, determining cell behaviour. Deciphering genotype-phenotype relationships has been crucial to understand normal and disease biology. Analysis of molecular pathways has provided an invaluable tool to such understanding; however, it does typically not consider the physical microenvironment, which is a key determinant of phenotype.\n\nIn this study, we present a novel modelling framework that enables to study the link between genotype, signalling networks and cell behaviour in a 3D microenvironment. To achieve this we bring together Agent Based Modelling, a powerful computational modelling technique, and gene networks. This combination allows biological hypotheses to be tested in a controlled stepwise fashion, and it lends itself naturally to model a heterogeneous population of cells acting and evolving in a dynamic microenvironment, which is needed to predict the evolution of complex multi-cellular dynamics. Importantly, this enables modelling co-occurring intrinsic perturbations, such as mutations, and extrinsic perturbations, such as nutrients availability, and their interactions.\n\nUsing cancer as a model system, we illustrate the how this framework delivers a unique opportunity to identify determinants of single-cell behaviour, while uncovering emerging properties of multi-cellular growth.\n\nAvailability and ImplementationFreely available on the web at http://www.microc.org. Research Resource Identification Initiative ID (https://scicrunch.org/): SCR 016672

bioinformatics

Genome-wide post-transcriptional dysregulation by microRNAs in human asthma as revealed by Frac-seq

MicroRNAs are small non-coding RNAs that inhibit gene expression post-transcriptionally, implicated in virtually all biological processes. Although the effect of individual microRNAs is generally studied, the genome-wide role of multiple microRNAs is less investigated. We assessed paired genome-wide expression of microRNAs with total (cytoplasmic) and translational (polyribosome-bound) mRNA levels employing Frac-seq in human primary bronchoepithelium from healthy controls and severe asthmatics. Severe asthma is a chronic inflammatory disease of the airways characterized by poor response to therapy. We found genes (=all isoforms of a gene) and mRNA isoforms differentially expressed in asthma, with novel inflammatory and structural mechanisms disclosed solely by polyribosome-bound mRNAs. Gene expression (=all isoforms of a gene) and mRNA expression analysis revealed different molecular candidates and biological pathways, with differentially expressed polyribosome-bound and total mRNAs also showing little overlap. We reveal a hub of six dysregulated microRNAs accounting for [~]90% of all microRNA targeting, displaying preference for polyribosome-bound mRNAs. Transfection of this hub in healthy cells mimicked asthma characteristics. Our work demonstrates extensive post-transcriptional gene dysregulation in asthma, where microRNAs play a central role, illustrating the feasibility and importance of assessing post-transcriptional gene expression when investigating human disease.

systems biology

LncRNA-TUG1/EZH2 Axis Promotes Cell Proliferation, Migration And The EMT Phenotype Formation Through Sponging miR-382

Pancreatic carcinoma (PC) is the one of the most common and malignant cancer in the world. Despite many effort have been made in recent years, the survival rate of PC still remains unsatisfied. Therefore, investigating the mechanisms underlying the progression of PC might facilitate the development of novel treatments that improve patient prognosis. LncRNA Taurine Up-regulated Gene 1 (TUG1) was initially identified as a transcript up - regulated by taurine, siRNA - based depletion of TUG1 suppresses mouse retinal development, and the abnormal expression of TUG1 has been reported in many cancers. However, the biological role and molecular mechanism of TUG1 in pancreatic carcinoma (PC) still needs to be further investigated. In the current study, the expression of TUG1 in the PC cell lines and tissues was measured by quantitative real-time PCR (qRT-PCR), and loss-of-function and gain-of-function approaches were applied to investigate the function of TUG1 in PC cell. Online database analysis tools showed that miR-382 could interact with TUG1 and we found an inverse correlation between TUG1 and miR-382 in PC specimens. Moreover, dual luciferase reporter assay, RNA-binding protein immunoprecipitation (RIP) and applied biotin-avidin pulldown system further provide evidence that TUG1 directly targeted miR-382 by binding with microRNA binding site harboring in the TUG1 sequence. Furthermore, gene expression array analysis using clinical samples and RT-qPCR proposed that EZH2 was a target of miR-382 in PC. Collectively, these findings revealed that TUG1 functions as an oncogenic lncRNA that promotes tumor progression at least partially through function as an endogenous sponge by competing for miR-382 binding to regulate the miRNA target EZH2.

cancer biology

Principles of studying a cell - a non-boastful paper for all molecular biologists

Studies of a cell rely on either observational approaches or perturbational/genetic approaches to define the contribution of a gene to specific cellular traits. It is unclear, however, under what circumstances each of the two approaches can be most successful and when they are doomed to fail. By analyzing over 500 complex traits of the yeast Saccharomyces cerevisiae we show that the trait relatedness to fitness determines the performance of observational approaches. Specifically, in traits subject to strong natural selection, genes identified using observational approaches are often highly coordinated in expression, such that the gene-trait associations are readily recognizable; in sharp contrast, the lack of such coordination in traits subject to weak selection leads to no detectable activity-trait associations for any individual genes and thus the failure of observational approaches. We further show that genetic approaches can be successful when the genes responsible for coordinating the target genes of observational approaches are perturbed. However, because the system-level cellular responses to a random mutation affect more or less every gene and consequently every trait, most genetic effects convey no trait-specific functional information for understanding the traits, which is particularly true for traits subject to weak selection.\n\nSignificance statementCell research is nearly exclusively based on empirical data obtained through either observational approaches or perturbational/genetic approaches. It is, however, increasingly clear that an analytical framework able to guide the empirical strategies is necessary to drive the field further ahead. This study analyzes ~500 complex traits of the yeast Saccharomyces cerevisiae and reveals the organizing principles of a cell. Specifically, a cell can be viewed as a factory, with each trait being the product of a production line operated directly by workers who are supervised by managers. For a cellular trait produced by many workers, the coordination level of the workers determines the performance of observational approaches. Meanwhile, the coordination of workers is realized by managers that are recruited and/or maintained by natural selection. Thus, observational approaches are expected to fail for traits subject to little selection, and genetic approaches can be successful only when the managers of fitness-tightly-coupled traits are perturbed. The manager-worker architecture built by natural selection explains well the origins of global epistasis and ubiquitous genetic effects, two major issues confusing current genetics and molecular and cellular biology, providing a clear guideline on how to study a cell.

Systems Biology

Automated structure refinement of macromolecular assemblies from cryo-EM maps using Rosetta

Cryo-EM has revealed many challenging yet exciting macromolecular assemblies at near-atomic resolution (3-4.5[A]), providing biological phenomena with molecular descriptions. However, at these resolutions accurately positioning individual atoms remains challenging and may be error-prone. Manually refining thousands of amino acids - typical in a macromolecular assembly - is tedious and time-consuming. We present an automated method that can improve the atomic details in models manually built in near-atomic-resolution cryo-EM maps. Applying the method to three systems recently solved by cryo-EM, we are able to improve model geometry while maintaining or improving the fit-to-density. Backbone placement errors are automatically detected and corrected, and the refinement shows a large radius of convergence. The results demonstrate the method is amenable to structures with symmetry, of very large size, and containing RNA as well as covalently bound ligands. The method should streamline the cryo-EM structure determination process, providing accurate and unbiased atomic structure interpretation of such maps.

Biochemistry

Anchor negatively regulates BMP signaling to control Drosophila wing development

Summary statementThe novel gene anchor is the ortholog of vertebrate GPR155, which contributes to preventing wing disc tissue overgrowth and limiting the phosphorylation of Mad in presumptive veins during the pupal stage.\n\nG protein-coupled receptors play a particularly important function in many organisms. The novel Drosophila gene anchor is the ortholog of vertebrate GPR155, and its molecular function and biological process are not yet known, especially in wing development. Knocking down anchor resulted in increased wing size and extra and thickened veins. These abnormal wing phenotypes are similar to those observed in gain-of-function of BMP signaling experiments. We observed that the BMP signaling indicator p-Mad was significantly increased in anchor RNAi-induced wing discs in larvae and that it also abnormally accumulated in intervein regions in pupae. Furthermore, the expression of BMP signaling pathway target genes were examined using a lacZ reporter, and the results indicated that omb and sal were substantially increased in anchor knockdown wing discs. In a study of genetic interactions between Anchor and BMP signaling pathway, the broadened and ectopic vein tissues were rescued by knocking down BMP levels. The results suggested that the function of Anchor is to negatively regulate BMP signaling during wing development and vein formation, and that Anchor targets or works upstream of Dpp.

Developmental Biology

SourceData - a semantic platform for curating and searching figures

Here we present SourceData (http://sourcedata.embo.org), a platform that allows researchers and publishers to share scientific figures and, when available, the underlying source data in a way that is machine-readable and findable. SourceData is unique in its focus on the core of scientific evidence--data presented in figures--and its capability to make papers searchable based on their data content and hence directly couple data to improved discoverability. SourceData aims at establishing a selfreinforcing data ecosystem that bridges the conventional visual and narrative description of research findings with a machine-readable representation of data and hypotheses.\n\nIn molecular and cell biology, most of the data that result from hypothesis-driven research are exclusively available in the form of figures or tables in published papers. In spite of their importance for the understanding of biological processes and human di ...

Bioinformatics

Reconstructing the backbone of the Saccharomycotina yeast phylogeny using genome-scale data

Understanding the phylogenetic relationships among the yeasts of the subphylum Saccharomycotina is a prerequisite for understanding the evolution of their metabolisms and ecological lifestyles. In the last two decades, the use of rDNA and multi-locus data sets has greatly advanced our understanding of the yeast phylogeny, but many deep relationships remain unsupported. In contrast, phylogenomic analyses have involved relatively few taxa and lineages that were often selected with limited considerations for covering the breadth of yeast biodiversity. Here we used genome sequence data from 86 publicly available yeast genomes representing 9 of the 11 major lineages and 10 non-yeast fungal outgroups to generate a 1,233-gene, 96-taxon data matrix. Species phylogenies reconstructed using two different methods (concatenation and coalescence) and two data matrices (amino acids or the first two codon positions) yielded identical and highly supported relationships between the 9 major lineages. Aside from the lineage comprised by the family Pichiaceae, all other lineages were monophyletic. Most interrelationships among yeast species were robust across the two methods and data matrices. However, 8 of the 93 internodes conflicted between analyses or data sets, including the placements of: the clade defined by species that have reassigned the CUG codon to encode serine, instead of leucine; the clade defined by a whole genome duplication; and of Ascoidea rubescens. These phylogenomic analyses provide a robust roadmap for future comparative work across the yeast subphylum in the disciplines of taxonomy, molecular genetics, evolutionary biology, ecology, and biotechnology. To further this end, we have also provided a BLAST server to query the 86 Saccharomycotina genomes, which can be found at http://y1000plus.org/blast.

Evolutionary Biology

Experimental and Computational Investigation of the Structure of Peptide Monolayers on Gold Nanoparticles

The self-assembly and self-organization of small molecules at the surface of nanoparticles constitute a potential route towards the preparation of advanced protein-like nanosystems. However, their structural characterization, critical to the design of bio-nanomaterials with well-defined biophysical and biochemical properties, remains highly challenging. Here, a computational model for peptide-capped gold nanoparticles is developed using experimentally characterized CALNN-and CFGAILSS-capped gold nanoparticles as a benchmark. The structure of CALNN and CFGAILSS monolayers is investigated by both structural biology techniques and molecular dynamics simulations. The calculations reproduce the experimentally observed dependence of the monolayer secondary structure on peptide capping density and on nanoparticle size, thus giving us confidence in the model. Furthermore, the computational results reveal a number of new features of peptide-capped monolayers, including the importance of sulfur movement for the formation of secondary structure motifs, the presence of water close to the gold surface even in tightly packed peptide monolayers, and the existence of extended 2D parallel {beta}-sheet domains in CFGAILSS monolayers. The model developed here provides a predictive tool that may assist in the design of further bio-nanomaterials.

biochemistry

HBP: an integrative and flexible pipeline for the interaction analysis of Hi-C dataset

BackgroundThe spatial organization of interphase chromatin in the nucleus play an important role in gene expression regulation and function. With the rapid development of revolutionized chromosome conformation capture technology and its genome-wide derivatives such as Hi-C, investigation of the genome folding becomes more efficient and convenient. How to robustly deal with these massive datasets and infer accurate 3D model and within-nucleus compartmentalization of chromosomes becomes a new challenge.\n\nResultThe implemented pipeline HBP (Hi-C BED file analysis Pipeline) integrates existing pipelines focusing on individual steps of Hi-C data processing into an all-in-one package with adjustable parameters to infer the consensus 3D structure of genome from raw Hi-C sequencing data. Whats more, HBP could assign statistical confidence estimation for chromatin interactions, and clustering interaction loci according to enrichment tracks or topological structure automatically.\n\nConclusionThe freely available HBP is an optimized and flexible pipeline for analyzing the folding of whole chromosome and interactions between some specific sites from the Hi-C raw sequencing reads to the partially processed datasets. The other complex genetic and epigenetic datasets from public sources such as GWAS, ENCODE consortiums etc. will also easily be integrated into HBP, hence the final output results of HBP could provide a comprehensive in-depth understanding for the specific chromatin interactions, potential molecular mechanisms and biological significance. We believe that HBP is a reliable tool for the rapidly analysis of Hi-C data and will be very useful for a wide range of researchers, particularly those who lack of background in computational biology. HBP is freely accessible at https://github.com/hechao0407/HBP/blob/master/HBP_1.0.tar.gz.

bioinformatics