Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Systems Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

SCORPIUS improves trajectory inference and identifies novel modules in dendritic cell development

1Recent advances in RNA sequencing enable the generation of genome-wide expression data at the single-cell level, opening up new avenues for transcriptomics and systems biology. A new application of single-cell whole-transcriptomics is the unbiased ordering of cells according to their progression along a dynamic process of interest. We introduce SCORPIUS, a method which can effectively reconstruct an ordering of individual cells without any prior information about the dynamic process. Comprehensive evaluation using ten scRNA-seq datasets shows that SCORPIUS consistently outperforms state-of-the-art techniques. We used SCORPIUS to generate novel hypotheses regarding dendritic cell development, which were subsequently validated in vivo. This work enables data-driven investigation and characterization of dynamic processes and lays the foundation for objective benchmarking of future trajectory inference methods.

Bioinformatics

The E. coli molecular phenotype under different growth conditions

Modern systems biology requires extensive, carefully curated measurements of cellular components in response to different environmental conditions. While high-throughput methods have made transcriptomics and proteomics datasets widely accessible and relatively economical to generate, systematic measurements of both mRNA and protein abundances under a wide range of different conditions are still relatively rare. Here we present a detailed, genome-wide transcriptomics and proteomics dataset of E. coli grown under 34 different conditions. We manipulate concentrations of sodium and magnesium in the growth media, and we consider four different carbon sources glucose, gluconate, lactate, and glycerol. Moreover, samples are taken both in exponential and stationary phase, and we include two extensive time-courses, with multiple samples taken between 3 hours and 2 weeks. We find that exponential-phase samples systematically differ from stationary-phase samples, in particular at the level of mRNA. Regulatory responses to different carbon sources or salt stresses are more moderate, but we find numerous differentially expressed genes for growth on gluconate and under salt and magnesium stress. Our data set provides a rich resource for future computational modeling of E. coli gene regulation, transcription, and translation.

bioinformatics

Searching for the chromatin determinants of human hematopoiesis

Hematopoiesis is one of the best characterized biological systems but the connection between chromatin changes and lineage differentiation is not yet well understood. We have developed a bioinformatic workflow to generate a chromatin space that allows to classify forty-two human healthy blood epigenomes from the BLUEPRINT, NIH ROADMAP and ENCODE consortia by their cell type. This approach let us to distinguish different cells types based on their epigenomic profiles, thus recapitulating important aspects of human hematopoiesis. The analysis of the orthogonal dimension of the chromatin space identify 32,662 chromatin determinant regions (CDRs), genomic regions with different epigenetic characteristics between the cell types. Functional analysis revealed that these regions are linked with cell identities. The inclusion of leukemia epigenomes in the healthy hematological chromatin sample space gives us insights on the healthy cell types that are more epigenetically similar to the disease samples. Further analysis of tumoral epigenetic alterations in hematopoietic CDRs points to sets of genes that are tightly regulated in leukemic transformations and commonly mutated in other tumors. Our method provides an analytical approach to study the relationship between epigenomic changes and cell lineage differentiation. Method availability: https://github.com/david-juan/ChromDet

genomics

Transcriptomics reveals patterns of sexually dimorphic gene expression in an avian hypothalamic-pituitary-gonadal (HPG) axis

The hypothalamic-pituitary-gonadal (HPG) axis is a key biological system required for reproduction and associated sexual behaviors to occur. Each of these tissues, the hypothalamus in the brain, the pituitary gland, and the gonads (testes and ovaries), has specialized and sometimes sex-specific functions that promote, amongst other things, reproductive processes. To fully characterize the transcript community of each aspect of the HPG axis, as well as to uncover potentially important sex-specific differences, we characterized patterns of gene expression in sexually mature male and female rock dove (Columba livia) hypothalamus, pituitary, and gonads. We describe patterns of gene expression amongst tissues as well as patterns of sex-biased gene expression which may underlie some of the fundamental differences between male and female reproductive behavior. We report greater sex-biased differential expression in the pituitary (233 genes out of 15,102 genes) as compared to the hypothalamus (1 gene out of 15,102 genes), with multiple genes overexpressed in the male pituitary gland being related to locomotion, and multiple genes over-expressed in the female pituitary gland being related to reproduction, growth and development. These genes may be associated with facilitating the different roles the HPG system plays in sex-specific reproductive behavior, including life-history strategies characterized by short-term payoffs in males (i.e. locomotion) and longer-term payoffs in females (i.e. development and reproduction). In addition, we report novel patterns of sex-biased expression in genes involved in reproduction-associated processes. Gonadotropin-releasing hormone, Progesterone, and Androgen receptor were upregulated in the female pituitary as compared to the male pituitary. We also found greater expression of Prolactin in the female pituitary as compared to the male pituitary, but it was more highly expressed in the male hypothalamus as compared to the female hypothalamus. The Prolactin receptor was also more highly expressed in the male hypothalamus and pituitary gland as compared to corresponding female tissues. Finally, we discovered greater expression of Arginine Vasopressin Receptor 1A in the male pituitary as compared to the female pituitary. Using a more global analysis, we discovered other interesting sex-biased patterns in genes not originally targeted for investigation. These genes include like Betacellulin (BTC), which was upregulated in the female pituitary and compared to males, and Ecto-NOX Disulfide-Thiol Exchanger 1 (ENOX1), which was upregulated in the male pituitary as compared to females. These genes may play important, though currently unknown, roles in reproductive physiology and behavior. In conclusion, we reveal patterns of tissue specific and sexually dimorphic gene expression in the HPG axis, highlighting the need for sex parity in transcriptomic studies and providing new lines of investigation of the mechanisms of reproductive function.

neuroscience

Flowtrace: simple visualization of coherent structures in biological fluid flows

1We present a simple, intuitive algorithm for visualizing time-varying flow fields that can reveal complex flow structures with minimal user intervention. We apply this technique to a variety of biological systems, including the swimming currents of invertebrates and the collective motion of swarms of insects. We compare our results to more experimentally-diffcult and mathematically-sophisticated techniques for identifying patterns in fluid flows, and suggest that our tool represents an essential \"middle ground\" allowing experimentalists to easily determine whether a system exhibits interesting flow patterns and coherent structures without the need to resort to more intensive techniques. In addition to being informative, the visualizations generated by our tool are often striking and elegant, illustrating coherent structures directly from videos without the need for computational overlays. Our tool is available as fully-documented open-source code available for MATLAB, Python, or ImageJ at www.flowtrace.org.

biophysics

Infectious Disease Dynamics Inferred from Genetic Data via Sequential Monte Carlo

Genetic sequences from pathogens can provide information about infectious disease dynamics that may supplement or replace information from other epidemiological observations. Currently available methods first estimate phylogenetic trees from sequence data, then estimate a transmission model conditional on these phylogenies. Outside limited classes of models, existing methods are unable to enforce logical consistency between the model of transmission and that underlying the phylogenetic reconstruction. Such conflicts in assumptions can lead to bias in the resulting inferences. Here, we develop a general, statistically efficient, plug-and-play method to jointly estimate both disease transmission and phylogeny using genetic data and, if desired, other epidemiological observations. This method explicitly connects the model of transmission and the model of phylogeny so as to avoid the aforementioned inconsistency. We demonstrate the feasibility of our approach through simulation and apply it to estimate stage-specific infectiousness in a subepidemic of HIV in Detroit, Michigan. In a supplement, we prove that our approach is a valid sequential Monte Carlo algorithm. While we focus on how these methods may be applied to population-level models of infectious disease, their scope is more general. These methods may be applied in other biological systems where one seeks to infer population dynamics from genetic sequences, and they may also find application for evolutionary models with phenotypic rather than genotypic data.

epidemiology

Visual motion computation in recurrent neural networks

Populations of neurons in primary visual cortex (V1) transform direct thalamic inputs into a cortical representation which acquires new spatio-temporal properties. One of these properties, motion selectivity, has not been strongly tied to putative neural mechanisms, and its origins remain poorly understood. Here we propose that motion selectivity is acquired through the recurrent mechanisms of a network of strongly connected neurons. We first show that a bank of V1 spatiotemporal receptive fields can be generated accurately by a network which receives only instantaneous inputs from the retina. The temporal structure of the receptive fields is generated by the long timescale dynamics associated with the high magnitude eigenvalues of the recurrent connectivity matrix. When these eigenvalues have complex parts, they generate receptive fields that are inseparable in time and space, such as those tuned to motion direction. We also show that the recurrent connectivity patterns can be learnt directly from the statistics of natural movies using a temporally-asymmetric Hebbian learning rule. Probed with drifting grating stimuli and moving bars, neurons in the model show patterns of responses analogous to those of direction-selective simple cells in primary visual cortex. These computations are enabled by a specific pattern of recurrent connections, that can be tested by combining connectome reconstructions with functional recordings.*\n\nAuthor summaryDynamic visual scenes provide our eyes with enormous quantities of visual information, particularly when the visual scene changes rapidly. Even at modest moving speeds, individual small objects quickly change their location causing single points in the scene to change their luminance equally fast. Furthermore, our own movements through the world add to the velocities of objects relative to our retinas, further increasing the speed at which visual inputs change. How can a biological system process efficiently such vast amounts of information, while keeping track of objects in the scene? Here we formulate and analyze a solution that is enabled by the temporal dynamics of networks of neurons.

neuroscience

Genomic and Environmental Contributions to Chronic Diseases in Urban Populations

Uncovering the interaction between genomes and the environment is a principal challenge of modern genomics and preventive medicine. While theoretical models are well defined, little is known of the GxE interactions in humans. We used a system biology approach to comprehensively assess the interactions between 1.6 million environmental exposure data, health, and expression phenotypes, together with whole genome genetic variation, for [~]1000 individuals from a founder-population in Quebec. We reveal a substantial impact of the urbanization gradient on the transcriptome and clinical endophenotypes, overpowering that of genetic ancestry. In detail, air pollution impacts gene expression and pathways affecting cardio-metabolic and respiratory traits when controlling for genetic ancestry. Finally, we capture 34 clinically associated expression quantitative trait loci that interact with the environment (air pollution). Our findings demonstrate how the local environment directly affects chronic disease development, and that genetic variation, including rare variants, can modulate individuals response to environmental challenges.\n\nHighlightsO_LIFine scale environmental effects overpower those of ancestry on gene expression\nC_LIO_LIAir pollution (geographic and temporal) is associated with transcriptional response\nC_LIO_LIGene-by-environment interactions with air pollution include asthma associated loci\nC_LIO_LIInflammatory pathways and cardio-respiratory clinical traits are among those affected\nC_LI

genomics

Dynamic species classification of microorganisms across time, abiotic and biotic environments - a sliding window approach

1. Technological advances have greatly simplified to take and analyze digital images and videos, and ecologists increasingly use these techniques for trait, behavioral and taxonomic analyses. The development of techniques to automate biological measurements from the environment opens up new possibilities to infer species numbers, observe presence/absence patterns and recognize individuals based on audio-visual information.\n\n2. Streams of quantitative data, such as temporal species abundances, are processed by machine learning (ML) algorithms into meaningful information. Machine learning approaches learn to distinguish classes (e.g., species) from observed quantitative features (phenotypes), and in-turn predict the distinguished classes in subsequent observations. However, in biological systems, the environment changes, often driving phenotypic changes in behaviour and morphology.\n\n3. Here we describe a framework for classifying species under dynamic biotic and abiotic conditions using a novel sliding window approach. We train a random forest classifier on subsets of the data, covering restricted temporal, biotic and abiotic ranges (i.e. windows). We test our approach by applying the classification framework to experimental microbial communities where results were validated against manual classification. Individuals from one to six ciliate species were monitored over hundreds of generations in dozens of different species combinations and over a temperature gradient. We describe the steps of our classification pipeline and systematically explore the effects of the abiotic and biotic environments as well as temporal effects on classification success.\n\n4. Differences in biotic and abiotic conditions caused simplistic classification approaches to be unsuccessful. In contrast, the sliding window approach allowed classification to be highly successful, because phenotypic differences driven by environmental change could be captured in the learning algorithm. Importantly, automatic classification showed comparable success compared to manual identifications.\n\n5. Our framework allows for reliable classification even in dynamic environmental contexts, and may help to improve long-term monitoring of species from environmental samples. It therefore has application in disciplines with automatic enumeration and phenotyping of organisms such as eco-toxicology, ecology and evolutionary ecology, and broad-scale environmental monitoring.

ecology

Short term changes in the proteome of human cerebral organoids induced by 5-methoxy-N,N-dimethyltryptamine

Dimethyltryptamines are hallucinogenic serotonin-like molecules present in traditional Amerindian medicine (e.g. Ayahuasca, Virola) recently associated with cognitive gains, antidepressant effects and changes in brain areas related to attention, self-referential thought, and internal mentation. Historical and technical restrictions impaired understanding how such substances impact human brain metabolism. Here we used shotgun mass spectrometry to explore proteomic differences induced by dimethyltryptamine (5-methoxy-N, N-dimethyltryptamine, 5-MeO-DMT) on human cerebral organoids. Out of the 6,728 identified proteins, 934 were found differentially expressed in 5-MeO-DMT-treated cerebral organoids. In silico systems biology analyses support 5-MeO-DMTs anti-inflammatory effects and reveal a modulation of proteins associated with the formation of dendritic spines, including proteins involved in cellular protrusion formation, microtubule dynamics and cytoskeletal reorganization. Proteins involved in long-term potentiation were modulated in a complex manner, with significant increases in the levels of NMDAR, CaMKII and CREB, but a reduction of PKA and PKC levels. These results offer possible mechanistic insights into the neuropsychological changes caused by the ingestion of substances rich in dimethyltryptamines.

neuroscience

Multi-scale Bayesian modeling of cryo-electron microscopy density maps

SummaryCryo-electron microscopy (cryo-EM) has become a mainstream technique for determining the structures of complex biological systems. However, accurate integrative structural modeling has been hampered by the challenges in objectively weighing cryo-EM data against other sources of information due to the presence of random and systematic errors, as well as correlations, in the data. To address these challenges, we introduce a Bayesian scoring function that efficiently and accurately ranks alternative structural models of a macromolecular system based on their consistency with a cryo-EM density map and other experimental and prior information. The accuracy of this approach is benchmarked using complexes of known structure and illustrated in three applications: the structural determination of the GroEL/GroES, RNA polymerase II, and exosome complexes. The approach is implemented in the open-source Integrative Modeling Platform (http://integrativemodeling.org), thus enabling integrative structure determination by combining cryo-EM data with other sources of information.\n\nHighlightsO_LIWe present a modeling approach to integrate cryo-EM data with other sources of information\nC_LIO_LIWe benchmark our approach using synthetic data on 21 complexes of known structure\nC_LIO_LIWe apply our approach to the GroEL/GroES, RNA polymerase II, and exosome complexes\nC_LI

bioinformatics

Uncovering Robust Patterns of MicroRNA Co-Expression across Cancers using Bayesian Relevance Networks

Co-expression networks have long been used as a tool for investigating the molecular circuitry governing biological systems. However, most algorithms for constructing co-expression networks were developed in the microarray era, before high-throughput sequencing--with its unique statistical properties--became the norm for expression measurement. Here we develop Bayesian Relevance Networks, an algorithm that uses Bayesian reasoning about expression levels to account for the differing levels of uncertainty in expression measurements between highly- and lowly-expressed entities, and between samples with different sequencing depths. It combines data from groups of samples (e.g., replicates) to estimate group expression levels and confidence ranges. It then computes uncertainty-moderated estimates of cross-group correlations between entities, and uses permutation testing to assess their statistical significance. Using large scale miRNA data from The Cancer Genome Atlas, we show that our Bayesian update of the classical Relevance Networks algorithm provides improved reproducibility in co-expression estimates and lower false discovery rates in the resulting co-expression networks. Software is available at www.perkinslab.ca/Software.html.

bioinformatics

Long-Term Imaging of Cellular Forces with High Precision by Elastic Resonator Interference Stress Microscopy

Cellular forces are crucial for many biological processes but current methods to image them have limitations with respect to online analysis, resolution and throughput. Here, we present a robust approach to measure mechanical cell-substrate interactions in diverse biological systems by interferometrically detecting deformations of an elastic micro-cavity. Elastic Resonator Interference Stress Microscopy (ERISM) yields stress maps with exceptional precision and large dynamic range (2 nm displacement resolution over a >1 m range, translating into 1 pN force sensitivity). This enables investigation of minute vertical stresses (<1 Pa) involved in podosome protrusion, protein specific cell-substrate interaction and amoeboid migration through spatial confinement in real time. ERISM requires no zero-force reference and avoids phototoxic effects, which facilitates force monitoring over multiple days and at high frame rates and eliminates the need to detach cells after measurements. This allows observation of slow processes like differentiation and further investigation of cells, e.g. by immunostaining.

biophysics

xMWAS: an R package for data-driven integration and differential network analysis

SummaryIntegrative omics is a central component of most systems biology studies. Computational methods are required for extracting meaningful relationships across different omics layers. Various tools have been developed to facilitate integration of paired heterogenous omics data; however most existing tools allow integration of only two omics datasets. Further-more, existing data integration tools do not incorporate additional steps of identifying sub-networks or communities of highly connected entities and evaluating the topology of the integrative network under different conditions. Here we present xMWAS, an R package for data integration, network visualization, clustering, differential network analysis of data from biochemical and phenotypic assays, and two or more omics platforms.\n\nAvailabilityhttps://sourceforge.net/projects/xmwas/\n\nContactkuppal2@emory.edu

bioinformatics

piClusterBusteR: Software For Automated Classification And Characterization Of piRNA Cluster Loci

BackgroundPiwi-interacting RNAs (piRNAs) are sRNAs that have a distinct biogenesis and molecular function from siRNAs and miRNAs. The piRNA pathway is well-conserved and shown to play an important role in the regulatory capacity of germline cells in Metazoans. Significant subsets of piRNAs are generated from discrete genomic loci referred to as piRNA clusters. Given that the contents of piRNA clusters dictate the target specificity of primary piRNAs, and therefore the generation of secondary piRNAs, they are of great significance when considering transcriptional and post-transcriptional regulation on a genomic scale. A quantitative comparison of top piRNA cluster composition can provide further insight into piRNA cluster biogenesis and function.\n\nResultsWe have developed software for general use, piClusterBusteR, which performs nested annotation of piRNA cluster contents to ensure high-quality characterization, provides a quantitative representation of piRNA cluster composition by feature, and makes available annotated and unannotated piRNA cluster sequences that can be utilized for downstream analysis. The data necessary to run piClusterBusteR and the skills necessary to execute this software on any species of interest are not overly burdensome for biological researchers.\n\npiClusterBusteR has been utilized to compare the composition of top piRNA generating loci amongst 13 Metazoan species. Characterization and quantification of cluster composition allows for comparison within piRNA clusters of the same species and between piRNA clusters of different species.\n\nConclusionsWe have developed a tool that accurately, automatically, and efficiently describes the contents of piRNA clusters in any biological system that utilizes the piRNA pathway. The results from piClusterBusteR have provided an in-depth description and comparison of the architecture of top piRNA clusters within and between 13 species, as well as a description of annotated and unannotated sequences from top piRNA cluster loci in these Metazoans.\n\npiClusterBusteR is available for download on GitHub: https://github.com/pschreiner/piClusterBuster

bioinformatics

Automated Recommendation Of Metabolite Substructures From Mass Spectra Using Frequent Pattern Mining

Despite the increasing importance of non-targeted metabolomics to answer various life science questions, extracting biochemically relevant information from metabolomics spectral data is still an incompletely solved problem. Most computational tools to identify tandem mass spectra focus on a limited set of molecules of interest. However, such tools are typically constrained by the availability of reference spectra or molecular databases, limiting their applicability to identify unknown metabolites. In contrast, recent advances in the field illustrate the possibility to expose the underlying biochemistry without relying on metabolite identification, in particular via substructure prediction. We describe an automated method for substructure recommendation motivated by association rule mining. Our framework captures potential relationships between spectral features and substructures learned from public spectral libraries. These associations are used to recommend substructures for any unknown mass spectrum. Our method does not require any predefined metabolite candidates, and therefore it can be used for the partial identification of unknown unknowns. The method is called MESSAR (MEtabolite SubStructure Auto-Recommender) and is implemented in a free online web service available at messar.biodatamining.be.\n\nAuthor SummaryMass spectrometry is one of most used techniques to detect and identify metabolites. However, learning metabolite structures directly from mass spectrometry data has always been a challenging task. Thousands of mass spectra from various biological systems still remain unanalyzed simply because no current bioinformatic tools are able to generate structural hypotheses. By manually studying mass spectra of standard compounds, chemists discovered that metabolites that share common substructures can also share spectral features. As data scientists, we believe that such relationships can be unraveled from massive structure and spectra data by machine learning. In this study, we adapted \"association rule mining\", traditionally used in market basket analysis, to structural and spectral data, allowing us to investigate all spectral features - metabolite substructures relationships. We further collected all statistically sound relationships into a database and used them to assign substructral hypotheses to unexplored spectra. We named our approach MESSAR, MEtabolite SubStructure Auto-Recommender, available to the metabolomics and mass spectrometry community as a free and open web service.

bioinformatics

A Systematic Analysis Of Atomic Protein-Ligand Interactions In The PDB

As the protein databank (PDB) recently passed the cap of 123,456 structures, it stands more than ever as an important resource not only to analyze structural features of specific biological systems, but also to study the prevalence of structural patterns observed in a large body of unrelated structures, that may reflect rules governing protein folding or molecular recognition. Here, we compiled a list of 11,016 unique structures of small-molecule ligands bound to proteins - 6,444 of which have experimental binding affinity - representing 750,873 protein-ligand atomic interactions, and analyzed the frequency, geometry and impact of each interaction type. We find that hydrophobic interactions are generally enriched in high-efficiency ligands, but polar interactions are over-represented in fragment inhibitors. While most observations extracted from the PDB will be familiar to seasoned medicinal chemists, less expected findings, such as the high number of C-H ...O hydrogen bonds or the relatively frequent amide-{pi} stacking between the backbone amide of proteins and aromatic rings of ligands, uncover underused ligand design strategies.

biophysics

Protein Features Identification For Machine Learning-Based Prediction Of Protein-Protein Interactions

The long awaited challenge of post-genomic era and systems biology research is computational prediction of protein-protein interactions (PPIs) that ultimately lead to protein functions prediction. The important research questions is how protein complexes with known sequence and structure be used to identify and classify protein binding sites, and how to infer knowledge from these classification such as predicting PPIs of proteins with unknown sequence and structure. Several machine learning techniques have been applied for the prediction of PPIs, but the accuracy of their prediction wholly depends on the number of features being used for training. In this paper, we have performed a survey of protein features used for the prediction of PPIs. The open research challenges and opportunities in the area have also been discussed.

bioinformatics