Search bioRxivSearch

Biology subjects

Richard Bonneau

Publications and source records attributed to Richard Bonneau.

7 recordsLinked to original sources

Temporal probabilistic modeling of bacterial compositions derived from 16S rRNA sequencing

The number of microbial and metagenomic studies has increased drastically due to advance-ments in next-generation sequencing-based measurement techniques. Statistical analysis and the validity of conclusions drawn from (time series) 16S rRNA and other metagenomic sequencing data is hampered by the presence of significant amount of noise and missing data (sampling zeros). Accounting uncertainty in microbiome data is often challenging due to the difficulty of obtaining biological replicates. Additionally, the compositional nature of current amplicon and metagenomic data differs from many other biological data types adding another challenge to the data analysis.\n\nTo address these challenges in human microbiome research, we introduce a novel probabilistic approach to explicitly model overdispersion and sampling zeros by considering the temporal correlation between nearby time points using Gaussian Processes. The proposed Temporal Gaussian Process Model for Compositional Data Analysis (TGP-CODA) shows superior modeling performance compared to commonly used Dirichlet-multinomial, multinomial, and non-parametric regression models on real and synthetic data. We demonstrate that the nonreplicative nature of human gut microbiota studies can be partially overcome by our method with proper experimental design of dense temporal sampling. We also show that different modeling ap-proaches have a strong impact on ecological interpretation of the data, such as stationarity, persistence, and environmental noise models.\n\nA Stan implementation of the proposed method is available under MIT license at https://github.com/tare/GPMicrobiome.

Ecology

Biophysically motivated regulatory network inference: progress and prospects

Via a confluence of genomic technology and computational developments the possibility of network inference methods that automatically learn large comprehensive models of cellular regulation is closer than ever. This perspective will focus on enumerating the elements of computational strategies that, when coupled to appropriate experimental designs, can lead to accurate large-scale models of chromatin-state and transcriptional regulatory structure and dynamics. We highlight four research questions that require further investigation in order to make progress in network inference: using overall constraints on network structure like sparsity, use of informative priors and data integration to constrain individual model parameters, estimation of latent regulatory factor activity under varying cell conditions, and new methods for learning and modeling regulatory factor interactions. We conclude that methods combining advances in these four categories of required effort with new genomic technologies will result in biophysically motivated dynamic genome-wide regulatory network models for several of the best studied organisms and cell types.

Bioinformatics

Fused regression for multi-source gene regulatory network inference

Understanding gene regulatory networks is critical to understanding cellular differentiation and response to external stimuli. Methods for global network inference have been developed and applied to a variety of species. Most approaches consider the problem of network inference independently in each species, despite evidence that gene regulation can be conserved even in distantly related species. Further, network inference is often confined to single data-types (single platforms) and single cell types. We introduce a method for multi-source network inference that allows simultaneous estimation of gene regulatory networks in multiple species or biological processes through the introduction of priors based on known gene relationships such as orthology incorporated using fused regression. This approach improves network inference performance even when orthology mapping and conservation are incomplete. We refine this method by presenting an algorithm that extracts the true conserved subnetwork from a larger set of potentially conserved interactions and demonstrate the utility of our method in cross species network inference. Last, we demonstrate our methods utility in learning from data collected on different experimental platforms.

Bioinformatics

Environmental gene regulatory influence networks in rice (Oryza sativa): response to water deficit, high temperature and agricultural environments

Environmental Gene Regulatory Influence Networks (EGRINs) coordinate the timing and rate of gene expression in response to environmental and developmental signals. EGRINs encompass many layers of regulation, which culminate in changes in the level of accumulated transcripts. Here we infer EGRINs for the response of five tropical Asian rice cultivars to high temperatures, water deficit, and agricultural field conditions, by systematically integrating time series transcriptome data (720 RNA-seq libraries), patterns of nucleosome-free chromatin (18 ATAC-seq libraries), and the occurrence of known cis-regulatory elements. First, we identify 5,447 putative target genes for 445 transcription factors (TFs) by connecting TFs with genes with known cis-regulatory motifs in nucleosome-free chromatin regions proximal to transcriptional start sites (TSS) of genes. We then use network component analysis to estimate the regulatory activity for these TFs from the expression of these putative target genes. Finally, we inferred an EGRIN using the estimated TFA as the regulator. The EGRIN included regulatory interactions between 4,052 target genes regulated by 113 TFs. We resolved distinct regulatory roles for members of a large TF family, including a putative regulatory connection between abiotic stress and the circadian clock, as well as specific regulatory functions for TFs in the drought response. TFA estimation using network component analysis is an effective way of incorporating multiple genome-scale measurements into network inference and that supplementing data from controlled experimental conditions with data from outdoor field conditions increases the resolution for EGRIN inference.

Systems Biology

4C-ker: A method to reproducibly identify genome-wide interactions captured by 4C-Seq experiments

4C-Seq has proven to be a powerful technique to identify genome-wide interactions with a single locus of interest (or \"bait\") that can be important for gene regulation. However, analysis of 4C-Seq data is complicated by the many biases inherent to the technique. An important consideration when dealing with 4C-Seq data is the differences in resolution of signal across the genome that result from differences in 3D distance separation from the bait. This leads to the highest signal in the region immediately surrounding the bait and increasingly lower signals in far-cis and trans. Another important aspect of 4C-Seq experiments is the resolution, which is greatly influenced by the choice of restriction enzyme and the frequency at which it can cut the genome. Thus, it is important that a 4C-Seq analysis method is flexible enough to analyze data generated using different enzymes and to identify interactions across the entire genome. Current methods for 4C-Seq analysis only identify interactions in regions near the bait or in regions located in far-cis and trans, but no method comprehensively analyzes 4C signals of different length scales. In addition, some methods also fail in experiments where chromatin fragments are generated using frequent cutter restriction enzymes. Here, we describe 4C-ker, a Hidden-Markov Model based pipeline that identifies regions throughout the genome that interact with the 4C bait locus. In addition we incorporate methods for the identification of differential interactions in multiple 4C-seq datasets collected from different genotypes or experimental conditions. Adaptive window sizes are used to correct for differences in signal coverage in near-bait regions, far-cis and trans chromosomes. Using several datasets, we demonstrate that 4C-ker outperforms all existing 4C-Seq pipelines in its ability to reproducibly identify interaction domains at all genomic ranges with different resolution enzymes.\n\nAUTHORS SUMMARYCircularized chromosome conformation capture, or 4C-Seq is a technique developed to identify regions of the genome that are in close spatial proximity to a single locus of interest ( bait). This technique is used to detect regulatory interactions between promoters and enhancers and to characterize the nuclear environment of different regions within and across different cell types. So far, existing methods for 4C-Seq data analysis do not comprehensively identify interactions across the entire genome due to biases in the technique that are related to the decrease in 4C signal that results from increased 3D distance from the bait. To compensate for these weaknesses in existing methods we developed 4C-ker, a method that explicitly models these biases to improve the analysis of 4C-Seq to better understand the genome wide interaction profile of an individual locus.

Bioinformatics

Robust Classification of Protein Variation Using Structural Modeling and Large-Scale Data Integration

Existing methods for interpreting protein variation focus on annotating mutation pathogenicity rather than detailed interpretation of variant deleteriousness and frequently use only sequence-based or structure-based information. We present VIPUR, a computational framework that seamlessly integrates sequence analysis and structural modeling (using the Rosetta protein modeling suite) to identify and interpret deleterious protein variants. To train VIPUR, we collected 9,477 protein variants with known effects on protein function from multiple organisms and curated structural models for each variant from crystal structures and homology models. VIPUR can be applied to mutations in any organisms proteome with improved generalized accuracy (AUROC .83) and interpretability (AUPR .87) compared to other methods. We demonstrate that VIPUR s predictions of deleteriousness match the biological phenotypes in ClinVar and provide a clear ranking of prediction confidence. We use VIPUR to interpret known mutations associated with inflammation and diabetes, demonstrating the structural diversity of disrupted functional sites and improved interpretation of mutations associated with human diseases. Lastly we demonstrate VIPUR s ability to highlight candidate genes associated with human diseases by applying VIPUR to de novo variants associated with autism spectrum disorders.

Bioinformatics

The interdependent nature of multi-loci associations can be revealed by 4C-Seq

Use of low resolution single cell DNA FISH and population based high resolution chromosome conformation capture techniques have highlighted the importance of pairwise chromatin interactions in gene regulation. However, it is unlikely that these associations act in isolation of other interacting partners within the genome. Indeed, the influence of multi-loci interactions in gene control remains something of an enigma as beyond low-resolution DNA FISH we do not have the appropriate tools to analyze these. Here we present a method that uses standard 4C-seq data to identify multi-loci interactions from the same cell. We demonstrate the feasibility of our method using 4C-seq data sets that identify known pairwise interactions involving the Tcrb and Igk antigen receptor enhancers, in addition to novel tri-loci associations. We further show that enhancer deletions not only interfere with tri-loci interactions in which they participate, but they also disrupt pairwise interactions between other partner enhancers and this disruption is linked to a reduction in their transcriptional output. These findings underscore the functional importance of hubs and provide new insight into chromatin organization as a whole. Our method opens the door for studying multi-loci interactions and their impact on gene regulation in other biological settings.

Genomics