Search bioRxivSearch

Biology subjects

Bonneau, R.

Publications and source records attributed to Bonneau, R..

15 recordsLinked to original sources

Normalization methods for microbial abundance data strongly affect correlation estimates

Consistent estimation of associations in microbial genomic survey count data is fundamental to microbiome research. Technical limitations, including compositionality, low sample sizes, and technical variability, obstruct standard application of association measures and require data normalization prior to estimating associations. Here, we investigate the interplay between data normalization and microbial association estimation by a comprehensive analysis of statistical consistency. Leveraging the large sample size of the American Gut Project (AGP), we assess the consistency of the two prominent linear association estimators, correlation and proportionality, under different sample scenarios and data normalization schemes, including RNA-seq analysis work flows and log-ratio transformations. We show that shrinkage estimation, a standard technique in high-dimensional statistics, can universally improve the quality of association estimates for microbiome data. We find that large-scale association patterns in the AGP data can be grouped into five normalization-dependent classes. Using microbial association network construction and clustering as examples of exploratory data analysis, we show that variance-stabilizing and log-ratio approaches provide for the most consistent estimation of taxonomic and structural coherence. Taken together, the findings from our reproducible analysis workflow have important implications for microbiome studies in multiple stages of analysis, particularly when only small sample sizes are available.

bioinformatics

Spatiotemporal Dynamics of Molecular Pathology in Amyotrophic Lateral Sclerosis

Paralysis occurring in amyotrophic lateral sclerosis (ALS) results from denervation of skeletal muscle as a consequence of motor neuron degeneration. Interactions between motor neurons and glia contribute to motor neuron loss, but the spatiotemporal ordering of molecular events that drive these processes in intact spinal tissue remains poorly understood. Here, we use spatial transcriptomics to obtain gene expression measurements of mouse spinal cords over the course of disease, as well as of postmortem tissue from ALS patients, to characterize the underlying molecular mechanisms in ALS. We identify novel pathway dynamics, regional differences between microglia and astrocyte populations at early time-points, and discern perturbations in several transcriptional pathways shared between murine models of ALS and human postmortem spinal cords.\n\nOne Sentence SummaryAnalysis of the ALS spinal cord using Spatial Transcriptomics reveals spatiotemporal dynamics of disease driven gene regulation.

neuroscience

The impact of endogenous retroviruses on nuclear organization and gene expression

BackgroundThe organization of chromatin in the nucleus plays an essential role in gene regulation. When considering the mammalian genome it is important to take into account that about half of the DNA is comprised of transposable elements. Given their repetitive nature, reads associated with these elements are generally discarded or randomly distributed among elements of the same type in genome-wide analyses. Thus, it is challenging to identify the activities and properties of individual transposons. As a result, we only have a partial understanding of how transposons contribute to chromatin folding and how they impact gene regulation.\n\nResultsUsing adapted PCR and Capture-based chromosome conformation capture (3C) approaches, collectively called 4Tran, we take advantage of the repetitive nature of transposons to capture interactions from multiple copies of endogenous retrovirus (ERVs) in the human and mouse genomes. With 4Tran-PCR, reads are selectively mapped to unique regions in the genome. This enables the identification of TE interaction profiles for individual ERV families and integration events specific to particular genomes. With this approach we demonstrate that transposons engage in long-range intra-chromosomal interactions guided by the separation of chromosomes into A and B compartments as well as topologically associated domains (TADs). In contrast to 4Tran-PCR, Capture-4Tran can uniquely identify both ends of an interaction that involve retroviral repeat sequences, providing a powerful tool for uncovering the individual TE insertions that interact with, and potentially regulate target genes.\n\nConclusions4Tran provides new insight into the manner in which transposons contribute to chromosome architecture and identifies target genes that transposable elements can potentially control.

genomics

Leveraging chromatin accessibility for transcriptional regulatory network inference in T Helper 17 Cells

Transcriptional regulatory networks (TRNs) provide insight into cellular behavior by describing interactions between transcription factors (TFs) and their gene targets. The Assay for Transposase Accessible Chromatin (ATAC)-seq, coupled with transcription-factor motif analysis, provides indirect evidence of chromatin binding for hundreds of TFs genome-wide. Here, we propose methods for TRN inference in a mammalian setting, using ATAC-seq data to influence gene expression modeling. We rigorously test our methods in the context of T Helper Cell Type 17 (Th17) differentiation, generating new ATAC-seq data to complement existing Th17 genomic resources (plentiful gene expression data, TF knock-outs and ChIP-seq experiments). In this resource-rich mammalian setting, our extensive benchmarking provides quantitative, genome-scale evaluation of TRN inference combining ATAC-seq and RNA-seq data. We refine and extend our previous Th17 TRN, using our new TRN inference methods to integrate all Th17 data (gene expression, ATAC-seq, TF KO, ChIP-seq). We highlight new roles for individual TFs and groups of TFs (\"TF-TF modules\") in Th17 gene regulation. Given the popularity of ATAC-seq, which provides high-resolution with low sample input requirements, we anticipate that application of our methods will improve TRN inference in new mammalian systems, especially in vivo, for cells directly from humans and animal models.

systems biology

Multi-study inference of regulatory networks for more accurate models of gene regulation

Gene regulatory networks are composed of sub-networks that are often shared across biological processes, cell-types, and organisms. Leveraging multiple sources of information, such as publicly available gene expression datasets, could therefore be helpful when learning a network of interest. Integrating data across different studies, however, raises numerous technical concerns. Hence, a common approach in network inference, and broadly in genomics research, is to separately learn models from each dataset and combine the results. Individual models, however, often suffer from under-sampling, poor generalization and limited network recovery. In this study, we explore previous integration strategies, such as batch-correction and model ensembles, and introduce a new multitask learning approach for joint network inference across several datasets. Our method initially estimates the activities of transcription factors, and subsequently, infers the relevant network topology. As regulatory interactions are context-dependent, we estimate model coefficients as a combination of both dataset-specific and conserved components. In addition, adaptive penalties may be used to favor models that include interactions derived from multiple sources of prior knowledge including orthogonal genomics experiments. We evaluate generalization and network recovery using examples from Bacillus subtilis and Saccharomyces cerevisiae, and show that sharing information across models improves network reconstruction. Finally, we demonstrate robustness to both false positives in the prior information and heterogeneity among datasets.

systems biology

Towards region-specific propagation of protein functions

MotivationDue to the nature of experimental annotation, most protein function prediction methods operate at the protein-level, where functions are assigned to full-length proteins based on overall similarities. However, most proteins function by interacting with other proteins or molecules, and many functional associations should be limited to specific regions rather than the entire protein length. Most domain-centric function prediction methods depend on accurate domain family assignments to infer relationships between domains and functions, with regions that are unassigned to a known domain-family left out of functional evaluation. Given the abundance of residue-level annotations currently available, we present a function prediction methodology that automatically infers function labels of specific protein regions using protein-level annotations and multiple types of region-specific features.\n\nResultsWe apply this method to local features obtained from InterPro, UniProtKB and amino acid sequences and show that this method improves both the accuracy and region-specificity of protein function transfer and prediction by testing on both human and yeast proteomes. We compare region-level predictive performance of our method against that of a whole-protein baseline method using a held-out dataset of proteins with structurally-verified binding sites and also compare protein-level temporal holdout predictive performances to expand the variety and specificity of GO terms we could evaluate. Our results can also serve as a starting point to categorize GO terms into site-specific and whole-protein terms and select prediction methods for different classes of GO terms.\n\nAvailabilityThe code is freely available at: https://github.com/ek1203/region_spec_func_pred

bioinformatics

deepNF: Deep network fusion for protein function prediction

The prevalence of high-throughput experimental methods has resulted in an abundance of large-scale molecular and functional interaction networks. The connectivity of these networks provide a rich source of information for inferring functional annotations for genes and proteins. An important challenge has been to develop methods for combining these heterogeneous networks to extract useful protein feature representations for function prediction. Most of the existing approaches for network integration use shallow models that cannot capture complex and highly-nonlinear network structures. Thus, we propose deepNF, a network fusion method based on Multimodal Deep Autoencoders to extract high-level features of proteins from multiple heterogeneous interaction networks. We apply this method to combine STRING networks to construct a common low-dimensional representation containing high-level protein features. We use separate layers for different network types in the early stages of the multimodal autoencoder, later connecting all the layers into a single bottleneck layer from which we extract features to predict protein function. We compare the cross-validation and temporal holdout predictive performance of our method with state-of-the-art methods, including the recently proposed method Mashup. Our results show that our method outperforms previous methods for both human and yeast STRING networks. We also show substantial improvement in the performance of our method in predicting GO terms of varying type and specificity.\n\nAvailabilitydeepNF is freely available at: https://github.com/VGligorijevic/deepNF

systems biology

A novel domain assembly routine for creating full-length models of membrane proteins from known domain structures

Membrane proteins composed of soluble and membrane domains are often studied one domain at a time. However, to understand the biological function of entire protein systems and their interactions with each other and drugs, knowledge of full-length structures or models is required. Although few computational methods exist that could potentially be used to model full-length constructs of membrane proteins, none of these methods are perfectly suited for the problem at hand. Existing methods either require an interface or knowledge of the relative orientations of the domains, are not designed for domain assembly, and none of them are developed for membrane proteins. Here we describe the first domain assembly protocol specifically designed for membrane proteins that assembles intra- and extracellular soluble domains and the transmembrane domain into models of the full-length membrane protein. Our protocol does not require an interface between the domains and samples possible domain orientations based on backbone dihedrals in the flexible linker regions, created via fragment insertion, while keeping the transmembrane domain fixed in the membrane. Our method, mp_domain_assembly, implemented in RosettaMP samples domain orientations close to the native structure and is best used in conjunction with experimental data to reduce the conformational search space.

biophysics

c-Maf-dependent regulatory T cells mediate immunological tolerance to intestinal microbiota

Both microbial and host genetic factors contribute to the pathogenesis of autoimmune disease1-4. Accumulating evidence suggests that microbial species that potentiate chronic inflammation, as in inflammatory bowel disease (IBD), often also colonize healthy individuals. These microbes, including the Helicobacter species, have the propensity to induce autoreactive T cells and are collectively referred to as pathobionts4-8. However, an understanding of how such T cells are constrained in healthy individuals is lacking. Here we report that host tolerance to a potentially pathogenic bacterium, Helicobacter hepaticus (H. hepaticus), is mediated by induction of ROR{gamma}t+Foxp3+ regulatory T cells (iTreg) that selectively restrain pro-inflammatory TH17 cells and whose function is dependent on the transcription factor c-Maf. Whereas H. hepaticus colonization of wild-type mice promoted differentiation of ROR{gamma}t-expressing microbe-specific iTreg in the large intestine, in disease-susceptible IL-10-deficient animals there was instead expansion of colitogenic TH17 cells. Inactivation of c-Maf in the Treg compartment likewise impaired differentiation of bacteria-specific iTreg, resulting in accumulation of H. hepaticus-specific inflammatory TH17 cells and spontaneous colitis. In contrast, ROR{gamma}t inactivation in Treg only had a minor effect on bacterial-specific Treg-TH17 balance, and did not result in inflammation. Our results suggest that pathobiont-dependent IBD is a consequence of microbiota-reactive T cells that have escaped this c-Maf-dependent mechanism of iTreg-TH17 homeostasis.

immunology

The Rosetta all-atom energy function for macromolecular modeling and design

Over the past decade, the Rosetta biomolecular modeling suite has informed diverse biological questions and engineering challenges ranging from interpretation of low-resolution structural data to design of nanomaterials, protein therapeutics, and vaccines. Central to Rosettas success is the energy function: amodel parameterized from small molecule and X-ray crystal structure data used to approximate the energy associated with each biomolecule conformation. This paper describes the mathematical models and physical concepts that underlie the latest Rosetta energy function, beta_nov15. Applying these concepts,we explain how to use Rosetta energies to identify and analyze the features of biomolecular models.Finally, we discuss the latest advances in the energy function that extend capabilities from soluble proteins to also include membrane proteins, peptides containing non-canonical amino acids, carbohydrates, nucleic acids, and other macromolecules.

biophysics

Explicit Modeling of RNA Stability Improves Large-Scale Inference of Transcription Regulation

Inference of eukaryotic transcription regulatory networks remains challenging due to the large number of regu-lators, combinatorial interactions, and redundant pathways. Even in the model system Saccharomyces cerevisiae, inference has performed poorly. Most existing inference algorithms ignore crucial regulatory components, like RNA stability and post-transcriptional modulation of regulators. Here we demonstrate that explicitly modeling tran-scription factor activity and RNA half-lives during inference of a genome-wide transcription regulatory network in yeast not only advances prediction performance, but also produces new insights into gene-and condition-specific variation of RNA stability. We curated a high quality gold standard reference network that we use for priors on network structure and model validation. We incorporate variation of RNA half-lives into the Inferelator inference framework, and show improved performance over previously described algorithms and over implementations of the algorithm that do not model RNA degradation. We recapitulate known condition-and gene-specific trends in RNA half-lives, and make new predictions about RNA half-lives that are confirmed by experimental data.

biophysics

An Adaptive Geometric Search Algorithm for Macromolecular Scaffold Selection

A wide variety of protein and peptidomimetic design tasks require matching functional three-dimensional motifs to potential oligomeric scaffolds. Enzyme design, for example, aims to graft active-site patterns typically consisting of 3 to 15 residues onto new protein surfaces. Identifying suitable proteins capable of scaffolding such active-site engraftment requires costly searches to identify protein folds that can provide the correct positioning of side chains to host the desired active site. Other examples of biodesign tasks that require simpler fast exact geometric searches of potential side chain positioning include mimicking binding hotspots, design of metal binding clusters and the design of modular hydrogen binding networks for specificity. In these applications the speed and scaling of geometric search limits downstream design to small patterns. Here we present an adaptive algorithm to searching for side chain take-off angles compatible with an arbitrarily specified functional pattern that enjoys substantive performance improvements over previous methods. We demonstrate this method in both genetically encoded (protein) and synthetic (peptidomimetic) design scenarios. Examples of using this method with the Rosetta framework for protein design are provided but our implementation is compatible with multiple protein design frameworks and is freely available as a set of python scripts (https://github.com/JiangTian/adaptive-geometric-search-for-protein-design).

bioengineering

Computing structure-based lipid accessibility of membrane proteins with mp_lipid_acc in RosettaMP

BackgroundMembrane proteins are vastly underrepresented in structural databases, which has led to a lack of computational tools and the corresponding inappropriate use of tools designed for soluble proteins. For membrane proteins, lipid accessibility is an essential property. Even though programs are available for sequence-based prediction of lipid accessibility and structure-based identification of solvent-accessible surface area, the latter does not distinguish between water accessible and lipid accessible residues in membrane proteins.\n\nResultsHere we present mp_lipid_acc, the first method to identify lipid accessible residues from the protein structure, implemented in the RosettaMP framework and available as a webserver. Our method uses protein structures transformed in membrane coordinates, for instance from PDBTM or OPM databases, and a defined membrane thickness to classify lipid accessibility of residues. mp_lipid_acc is applicable to both -helical and {beta}-barrel membrane proteins of diverse architectures with or without water-filled pores and uses a concave hull algorithm for classification. We further provide a manually curated benchmark dataset, on which our method achieves prediction accuracies of 90%.\n\nConclusionWe present a novel tool to classify lipid accessibility from the protein structure, which is applicable to proteins of diverse architectures and achieves prediction accuracies of 90% on a manually curated database. mp_lipid_acc is part of the Rosetta software suite, available at www.rosettacommons.org. The webserver is available at http://rosie.graylab.jhu.edu/mp_lipid_acc/submit and the benchmark dataset is available at http://tinyurl.com/mp-lipid-acc-dataset.\n\nSupplementary informationSupplementary information is available at BMC Bioinformatics.

biophysics

Rotamer libraries for the high-resolution design of beta-amino acid foldamers

{beta}-amino acids offer attractive opportunities to develop biologically active peptidomimetics, either employed alone or in conjunction with natural -amino acids. Owing to their potential for unique conformational preferences that deviate considerably from -peptide geometries, {beta}-amino acids greatly expand the possible chemistries and physical properties available to polyamide foldamers. Complete in silico support for designing new molecules incorporating nonnatural amino acids typically requires representing their side chain conformations as sets of discrete rotamers for model refinement and sequence optimization. Such rotamer libraries are key components of several state of the art design frameworks. Here we report the development, incorporation in to the Rosetta macromolecular modeling suite, and validation of rotamer libraries for {beta}3-amino acids.

bioinformatics

Mocap: Large-scale inference of transcription factor binding sites from chromatin accessibility

Differential binding of transcription factors (TFs) at cis-regulatory loci drives the differentiation and function of diverse cellular lineages. Understanding the regulatory interactions that underlie cell fate decisions requires characterizing TF binding sites (TFBS) across multiple cell types and conditions. Techniques, e.g. ChIP-Seq can reveal genome-wide patterns of TF binding, but typically requires laborious and costly experiments for each TF-cell-type (TFCT) condition of interest. Chromosomal accessibility assays can connect accessible chromatin in one cell type to many TFs through sequence motif mapping. Such methods, however, rarely take into account that the genomic context preferred by each factor differs from TF to TF, and from cell type to cell type. To address the differences in TF behaviors, we developed Mocap, a method that integrates chromatin accessibility, motif scores, TF footprints, CpG/GC content, evolutionary conservation and other factors in an ensemble of TFCT-specific classifiers. We show that integration of genomic features, such as CpG islands improves TFBS prediction in some TFCT. Further, we describe a method for mapping new TFCT, for which no ChIP-seq data exists, onto our ensemble of classifiers and show that our cross-sample TFBS prediction method outperforms several previously described methods.

bioinformatics