Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Molecular Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,531 records · Page 85Linked to original sources

Photoacoustic molecular rulers based on DNA nanostructures

Molecular rulers that rely on the Forster resonance energy transfer (FRET) mechanism are widely used to investigate dynamic molecular processes that occur on the nanometer scale. However, the capabilities of these fluorescence molecular rulers are fundamentally limited to shallow imaging depths by light scattering in biological samples. Photoacoustic tomography (PAT) has recently emerged as a high resolution modality for in vivo imaging, coupling optical excitation with ultrasound detection. In this paper, we report the capability of PAT to probe distance-dependent FRET at centimeter depths. Using DNA nanotechnology we created several nanostructures with precisely positioned fluorophore-quencher pairs over a range of nanoscale separation distances. PAT of the DNA nanostructures showed distance-dependent photoacoustic signal generation and experimentally demonstrated the ability of PAT to reveal the FRET process deep within tissue mimicking phantoms. Further, we experimentally validated these DNA nanostructures as providing a novel and biocompatible strategy to augment the intrinsic photoacoustic signal generation capabilities of small molecule fluorescent dyes.

bioengineering

Consequences Of Natural Perturbations In The Human Plasma Proteome

Proteins are the primary functional units of biology and the direct targets of most drugs, yet there is limited knowledge of the genetic factors determining inter-individual variation in protein levels. Here we reveal the genetic architecture of the human plasma proteome, testing 10.6 million DNA variants against levels of 2,994 proteins in 3,301 individuals. We identify 1,927 genetic associations with 1,478 proteins, a 4-fold increase on existing knowledge, including trans associations for 1,104 proteins. To understand consequences of perturbations in plasma protein levels, we introduce an approach that links naturally occurring genetic variation with biological, disease, and drug databases. We provide insights into pathogenesis by uncovering the molecular effects of disease-associated variants. We identify causal roles for protein biomarkers in disease through Mendelian randomization analysis. Our results reveal new drug targets, opportunities for matching existing drugs with new disease indications, and potential safety concerns for drugs under development.

genomics

Topology and cooperative stability: the two master regulators of protein half-life in the cell

In a quest for finding additional structural constraints, apart from disordered segments, regulating protein half-life in the cell (and during evolution), here we recognize and assess the influence of native topology of biological proteins and their sequestration into multimeric complexes. Native topology acts as a molecular marker of proteins mechanical resistance and consequently captures their half-life variations on genome-scale, irrespective of the enormous sequence, structural and functional diversity of the proteins. Cooperative stability (slower degradation upon sequestration into complexes) is a master regulator of oligomeric protein half-life that involves at least three mechanisms. (i) Association with multiple complexes results longer protein half-life; (ii) hierarchy of complex self-assembly involves short-living proteins binding late in the assembly order and (iii) binding with larger buried surface area leads to slower subunit dissociation and thereby longer half-life. Altered half-lives of paralog proteins refer to their structural divergence and oligomerization with non-identical set of complexes.

biophysics

Transcriptome analysis of early stage of neurogenesis reveals regulatory gene network for preplate neuron differentiation and CR cell specification

Neurogenesis in the developing neocortex begins with the generation of the preplate, which consists of early born neurons including Cajal-Retzius (CR) cells and subplate neurons. Here, utilizing the Ebf2-EGFP transgenic mouse in which EGFP initially labels the preplate neurons then persists in CR cells, we reveal the dynamic transcriptome profiles of early neurogenesis and CR cell differentiation. At E15.5 when Ebf2-EGFP+ cells are mostly CR neurons, single-cell sequencing analysis of purified Ebf2-EGFP+ cells uncovers molecular heterogeneity in CR neurons, but without apparent clustering of cells with distinct regional origins. Along a pseudotemporal trajectory these cells are classified into three different developing states, revealing genetic cascades from early generic neuronal differentiation to late fate specification during the establishment of CR neuron identity and function. Further genome-wide RNA-seq and ChIP-seq analyses at multiple early neurogenic stages have revealed the temporal gene expression dynamics of early neurogenesis and distinct histone modification patterns in early differentiating neurons. We have also identified a new set of coding genes and lncRNAs involved in early neuronal differentiation and validated with functional assays In Vitro and In Vivo. Our findings shed light on the molecular mechanisms governing the early differentiation steps during cortical development, especially CR neuron biology, and help understand the developmental basis for cortical function and diseases.

neuroscience

DNA Methylation Network Estimation with Sparse Latent Gaussian Graphical Model

Inferring molecular interaction networks from genomics data is important for advancing our understanding of biological processes. Whereas considerable research effort has been placed on inferring such networks from gene expression data, network estimation from DNA methylation data has received very little attention due to the substantially higher dimensionality and complications with result interpretation for non-genic regions. To combat these challenges, we propose here an approach based on sparse latent Gaussian graphical model (SLGGM). The core idea is to perform network estimation on q latent variables as opposed to d CpG sites, with q<<d. To impose a correspondence between the latent variables and genes, we use the distance between CpG sites and transcription starting sites of the genes to generate a prior on the CpG sites latent class membership. We evaluate this approach on synthetic data, and show on real data that the gene network estimated from DNA methylation data significantly explains gene expression patterns in unseen datasets.

genomics

Decomposing cell identity for transfer learning across cellular measurements, platforms, tissues, and species.

New approaches are urgently needed to glean biological insights from the vast amounts of single cell RNA sequencing (scRNA-Seq) data now being generated. To this end, we propose that cell identity should map to a reduced set of factors which will describe both exclusive and shared biology of individual cells, and that the dimensions which contain these factors reflect biologically meaningful relationships across different platforms, tissues and species. To find a robust set of dependent factors in large-scale scRNA- Seq data, we developed a Bayesian non-negative matrix factorization (NMF) algorithm, scCoGAPS. Application of scCoGAPS to scRNA-Seq data obtained over the course of mouse retinal development identified gene expression signatures for factors associated with specific cell types and continuous biological processes. To test whether these signatures are shared across diverse cellular contexts, we developed projectR to map biologically disparate datasets into the factors learned by scCoGAPS. Because projecting these dimensions preserve relative distances between samples, biologically meaningful relationships/factors will stratify new data consistent with their underlying processes, allowing labels or information from one dataset to be used for annotation of the other--a machine learning concept called transfer learning. Using projectR, data from multiple datasets was used to annotate latent spaces and reveal novel parallels between developmental programs in other tissues, species and cellular assays. Using this approach we are able to transfer cell type and state designations across datasets to rapidly annotate cellular features in a new dataset without a priori knowledge of their type, identify a species-specific signature of microglial cells, and identify a previously undescribed subpopulation of neurosecretory cells within the lung. Together, these algorithms define biologically meaningful dimensions of cellular identity, state, and trajectories that persist across technologies, molecular features, and species.\n\nGRAPHICAL ABSTRACT\n\nO_FIG O_LINKSMALLFIG WIDTH=174 HEIGHT=200 SRC=\"FIGDIR/small/395004_ufig1.gif\" ALT=\"Figure 1\">\nView larger version (81K):\norg.highwire.dtl.DTLVardef@dd1c07org.highwire.dtl.DTLVardef@5b1109org.highwire.dtl.DTLVardef@bb6714org.highwire.dtl.DTLVardef@16c66f0_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics

Application of Atlas of Cancer Signalling Network in pre-clinical studies

Initiation and progression of cancer involve multiple molecular mechanisms. The knowledge on these mechanisms is expanding and should be converted into guidelines for tackling the disease. We discuss here formalization of biological knowledge into a comprehensive resource Atlas of Cancer Signalling Network (ACSN) and Google Maps-based tool NaviCell that supports map navigation. The application of maps for omics data visualisation in the context of signalling maps is possible using NaviCell Web Service module and NaviCom tool for generation of network-based molecular portraits of cancer using multi-level omics data. We review how these resources and tools are applied for cancer pre-clinical studies among others for rationalizing synergistic effect of drugs and designing complex disease stage-specific druggable interventions following structural analysis of the maps together with omics data. Modules and maps of ACSN as signatures of biological functions, can help in cancer data analysis and interpretation. In addition, they can also be used to find association between perturbations in particular molecular mechanisms to the risk of a specific cancer type development. These approaches and beyond help to study interplay between molecular mechanisms of cancer, deciphering how gene interactions govern hallmarks of cancer in specific context. We discuss a perspective to develop a flexible methodology and a pipeline to enable systematic omics data analysis in the context of signalling network maps, for stratifying patients and suggesting interventions points and drug repositioning in cancer and other human diseases.

systems biology

polyCluster: Defining Communities of Reconciled Cancer Subtypes with Biological and Prognostic Significance

To stratify cancer patients for most beneficial therapies, it is a priority to define robust molecular subtypes using clustering methods and \"big data\". If each of these methods produces different numbers of clusters for the same data, it is difficult to achieve an optimal solution. Here, we introduce \"polyCluster\", a tool that reconciles clusters identified by different methods into context-specific subtype \"communities\" using a hypergeometric test or a measure of relative proportion of common samples. The polycluster was tested using a breast cancer dataset, and latter using uveal melanoma datasets to identify novel subtype communities with significant metastasis-free prognostic differences. Available at: https://github.com/syspremed/polyClustR

genomics

Avoidance of toxic misfolding does not explain the sequence constrains of highly expressed proteins across organisms

The avoidance of cytotoxic effects associated with protein misfolding has been proposed as a dominant constraint on the sequence evolution and molecular clock of highly expressed proteins. Recently, Leuenberger et al. developed an elegant experimental approach to measure protein thermal stability at the proteome scale. The collected data allow us to rigorously test the predictions of the misfolding avoidance hypothesis that highly expressed proteins have evolved to be more stable, and that maintaining thermodynamic stability significantly constrains their evolution. Notably, careful re-analysis of the Leuenberger et al. data across four different organisms reveals no substantial correlation between protein stability and protein abundance. Therefore, the key predictions of the misfolding toxicity and related hypotheses are not supported by available empirical data. The data also suggest that, regardless of protein expression, protein stability does not substantially affect the protein molecular clock across organisms.

evolutionary biology

Positional information encoded in the dynamic differences between neighbouring oscillators during vertebrate segmentation.

How cells track their position during the segmentation of the vertebrate body remains elusive. For decades, this process has been interpreted according to the clock-and-wavefront model, where molecular oscillators set the frequency of somite formation while the positional information is encoded in signaling gradients. Recent experiments using ex vivo explants challenge this interpretation, suggesting that positional information is encoded in the properties of the oscillators. Here, we propose that positional information is encoded in the differential levels of neighboring oscillators. The differences gradually increase because both the oscillator amplitude and the period increase with time. When this difference exceeds a certain threshold, the segmentation program starts. Using this framework, we quantitatively fit experimental data from in vivo and ex vivo mouse segmentation, and propose mechanisms of somite scaling. Our results suggest a novel mechanism of spatial pattern formation based on the local interactions between dynamic molecular oscillators.

developmental biology

Fast calcium transients in neuronal spines is driven by extreme statistics

Extreme statistics describe the distribution of rare events that can define the timescales of transduction within cellular microdomains. We combine biophysical modeling and analysis of live-cell calcium imaging to explain the fast calcium transient in spines. We show that in the presence of a spine apparatus (SA), which is an extension of the smooth endoplasmic reticulum (ER), calcium transients during synaptic inputs rely on rare and extreme calcium ion trajectories. Using numerical simulations, we predicted the asymmetrical distributions of Ryanodine receptors and SERCA pumps that we confirmed experimentally. When calcium ions are released in the spine head, the fastest ions arriving at the base determine the transient timescale through a calcium-induced calcium release mechanism. In general, the fastest particles arriving at a small target are likely to be a generic mechanism that determines the timescale of molecular transduction in cellular neuroscience.\n\nSignificance statementIntrigued by fast calcium transients of few milliseconds in dendritic spines, we investigated its underlying biophysical mechanism. We show here that it is generated by the diffusion of the fastest calcium ions when the spine contains a Spine Apparatus, an extension of the endoplasmic reticulum. This timescale is modulated by the initial number of released calcium ions and the asymmetric distribution of its associated calcium release associated Ryanodyne receptors, present only at the base of a spine. This novel mechanism of calcium signaling that we have unraveled here is driven by the fastest particles. To conclude, the rate of arrival of the fastest particles (ions) to a small target receptor defines the timescale of activation instead of the classical forward rate of chemical reactions introduced by von Smoluchowski in 1916. Applying this new rate theory to transduction should refine our understanding of the biophysical mechanisms underlying molecular signaling.

cell biology

Transcriptome Analysis of Adult C. elegans Cells Reveals Tissue-specific Gene and Isoform Expression

The biology and behavior of adults differ substantially from those of developing animals, and cell-specific information is critical for deciphering the biology of multicellular animals. Thus, adult tissue-specific transcriptomic data are critical for understanding molecular mechanisms that control their phenotypes. We used adult cell-specific isolation to identify the transcriptomes of C. elegans four major tissues (or \"tissue-ome\"), identifying ubiquitously expressed and tissue-specific \"super-enriched\" genes. These data newly reveal the hypodermis metabolic character, suggest potential worm-human tissue orthologies, and identify tissue-specific changes in the Insulin/IGF-1 signaling pathway. Tissue-specific alternative splicing analysis identified a large set of collagen isoforms and a neuron-specific CREB isoform. Finally, we developed a machine learning-based prediction tool for 70 sub-tissue cell types, which we used to predict cellular expression differences in IIS/FOXO signaling, stage-specific TGF-b activity, and basal vs. memory-induced CREB transcription. Together, these data provide a rich resource for understanding the biology governing multicellular adult animals

genomics

Mergeomics: integration of diverse genomics resources to identify pathogenic perturbations to biological systems

Mergeomics is a computational pipeline (http://mergeomics.research.idre.ucla.edu/Download/Package/) that integrates multidimensional omics-disease associations, functional genomics, canonical pathways and gene-gene interaction networks to generate mechanistic hypotheses. It first identifies biological pathways and tissue-specific gene subnetworks that are perturbed by disease-associated molecular entities. The disease-associated subnetworks are then projected onto tissue-specific gene-gene interaction networks to identify local hubs as potential key drivers of pathological perturbations. The pipeline is modular and can be applied across species and platform boundaries, and uniquely conducts pathway/network level meta-analysis of multiple genomic studies of various data types. Application of Mergeomics to cholesterol datasets revealed novel regulators of cholesterol metabolism.

Systems Biology

Network-aware mutation clustering of cancer

The grouping of cancers across tissue boundaries is central to precision oncology, but remains a difficult problem. Here we present EPICC (Experimental Protein Interaction Clustering of Cancer), a novel technique to cluster cancer patients based on DNA mutation profile, that leverages knowledge of protein-protein interactions to reduce noise and amplify biological signal. We applied EPICC to data from The Cancer Genome Atlas (TCGA), and both recapitulated known cancer clusterings, and identified new cross-tissue cancer groups that may indicate novel cancer molecular subtypes. Investigation of EPICC clusters revealed new protein modules which were recurrently mutated across cancers, and indicate new avenues for research into cancer biology. EPICC leveraged the Vodafone DreamLab citizen science platform, and we provide our results as a resource for researchers to investigate the role of protein modules in cancer.

bioinformatics

RUNX3 regulates cell cycle-dependent chromatin dynamics by functioning as a pioneer factor of the restriction point

The cellular decision regarding whether to undergo proliferation or death is made at the restriction (R)-point, which is disrupted in nearly all tumors. The identity of the molecular mechanisms that govern the R-point decision is one of the fundamental issues in cell biology. We found that early after mitogenic stimulation, RUNX3 bound to its target loci, where it opened chromatin structure by sequential recruitment of Trithorax group proteins and cell-cycle regulators to drive cells to the R-point. Soon after, RUNX3 closed these loci by recruiting Polycomb repressor complexes, causing the cell to pass through the R-point toward S phase. If the RAS signal was constitutively activated, RUNX3 inhibited cell cycle progression by maintaining R-point-associated genes in an open structure. Our results identify RUNX3 as a pioneer factor for the R-point and reveal the molecular mechanisms by which appropriate chromatin modifiers are selectively recruited to target loci for appropriate R-point decisions.

cell biology

Connecting tumor genomics with therapeutics through multi-dimensional network modules

Recent efforts have catalogued genomic, transcriptomic, epigenetic and proteomic changes in tumors, but connecting these data with effective therapeutics remains a challenge. In contrast, cancer cell lines can model therapeutic responses but only partially reflect tumor biology. Bridging this gap requires new methods of data integration to identify a common set of pathways and molecular events. Using MAGNETIC, a new method to integrate molecular profiling data using functional networks, we identify 219 gene modules in TCGA breast cancers that capture recurrent alterations, reveal new roles for H3K27 tri-methylation and accurately quantitate various cell types within the tumor microenvironment. We show that a significant portion of gene expression and methylation in tumors is poorly reproduced in cell lines due to differences in biology and microenvironment and MAGNETIC identifies therapeutic biomarkers that are robust to these differences. This work addresses a fundamental challenge in pharmacogenomics that can only be overcome by the joint analysis of patient and cell line data.

bioinformatics

iOmicsPASS: a novel method for integration of multi-omics data over biological networks and discovery of predictive subnetworks

We developed iOmicsPASS, an intuitive method for network-based multi-omics data integration and detection of biological subnetworks for phenotype prediction. The method converts abundance measurements into co-expression scores of biological networks and uses a powerful phenotype prediction method adapted for network-wise analysis. Simulation studies show that the proposed data integration approach considerably improves the quality of predictions. We illustrate iOmicsPASS through the integration of quantitative multi-omics data using transcription factor regulatory network and protein-protein interaction network for cancer subtype prediction. Our analysis of breast cancer data identifies network signatures surrounding established markers of molecular subtypes. The analysis of colorectal cancer data highlights a protein interactome surrounding key proto-oncogenes as predictive features of subtypes, rendering them more biologically interpretable than the approaches integrating data without a priori relational information. However, the results indicate that current molecular subtyping is overly dependent on transcriptomic data and crude integrative analysis fails to account for molecular heterogeneity in other -omics data. The analysis also suggest that tumor subtypes are not mutually exclusive and future subtyping should therefore consider multiplicity in assignments.\n\nAvailability: https://github.com/cssblab/iOmicsPASS

systems biology

Exploring the phenotypic space and the evolutionary history of a natural mutation in Drosophila melanogaster

A major challenge of modern Biology is elucidating the functional consequences of natural mutations. While we have a good understanding of the effects of lab-induced mutations on the molecular- and organismal-level phenotypes, the study of natural mutations has lagged behind. In this work, we explore the phenotypic space and the evolutionary history of a previously identified adaptive transposable element insertion. We first combined several tests that capture different signatures of selection to show that there is evidence of positive selection in the regions flanking FBti0019386 insertion. We then explored several phenotypes related to known phenotypic effects of nearby genes, and having plausible connections to fitness variation in nature. We found that flies with FBti0019386 insertion had a shorter developmental time and were more sensitive to stress, which are likely to be the adaptive effect and the cost of selection of this mutation, respectively. Interestingly, these phenotypic effects are not consistent with a role of FBti0019386 in temperate adaptation as has been previously suggested. Indeed, a global analysis of the population frequency of FBti0019386 showed that clinal frequency patterns are found in North America and Australia but not in Europe. Finally, we showed that FBti0019386 is associated with down-regulation of sra most likely because it induces the formation of heterochromatin by recruiting HP1a protein. Overall, our integrative approach allowed us to shed light on the evolutionary history, the relevant fitness effects and the likely molecular mechanisms of an adaptive mutation and highlights the complexity of natural genetic variants.

Evolutionary Biology