Search bioRxivSearch

Biology subjects

Erik van Nimwegen

Publications and source records attributed to Erik van Nimwegen.

6 recordsLinked to original sources

Automated Incorporation of Pairwise Dependency in Transcription Factor Binding Site Prediction Using Dinucleotide Weight Tensors

Gene regulatory networks are ultimately encoded by the sequence-specific binding of (TFs) to short DNA segments. Although it is customary to represent the binding specificity of a TF by a position-specific weight matrix (PSWM), which assumes each position within a site contributes independently to the overall binding affinity, evidence has been accumulating that there can be significant dependencies between positions. Unfortunately, methodological challenges have so far hindered the development of a practical and generally-accepted extension of the PSWM model. On the one hand, simple models that only consider dependencies between nearest-neighbor positions are easy to use in practice, but fail to account for the distal dependencies that are observed in the data. On the other hand, models that allow for arbitrary dependencies are prone to overfitting, requiring regularization schemes that are difficult to use in practice for non-experts.\n\nHere we present a new regulatory motif model, called dinucleotide weight tensor (DWT), that incorporates arbitrary pairwise dependencies between positions in binding sites, rigorously from first principles, and free from tunable parameters. We demonstrate the power of the method on a large set of ChIP-seq data-sets, showing that DWTs outperform both PSWMs and motif models that only incorporate nearest-neighbor dependencies. We also demonstrate that DWTs outperform two previously proposed methods. Finally, we show that DWTs inferred from ChIP-seq data also outperform PSWMs on HT-SELEX data for the same TF, suggesting that DWTs capture inherent biophysical properties of the interactions between the DNA binding domains of TFs and their binding sites.\n\nWe make a suite of DWT tools available at dwt.unibas.ch, that allow users to automatically perform motif finding, i.e. the inference of DWT motifs from a set of sequences, binding site prediction with DWTs, and visualization of DWT dilogo motifs.\n\nAuthor SummaryGene regulatory networks are ultimately encoded in constellations of short binding sites in the DNA and RNA that are recognized by regulatory factors such as transcription factors (TFs). For several decades, computational analysis of regulatory networks has relied on a model of TF sequence-specificity, the position-specific weight-matrix (PSWM), that assumes different positions in a binding site contribute independently to the total binding energy of the TF. However, in recent years evidence has been accumulating that, at least for some TFs, this assumption does not hold. Here we present a new model for the sequence-specificity of TFs, the dinucleotide weight tensor (DWT), that takes arbitrary dependencies between positions in binding sites into account and show that it consistently outperforms PSWMs on high-throughput datasets on TF binding. Moreover, in contrast to previous approaches, DWTs are directly derived from first principles within a Bayesian framework, and contain no tunable parameters. This allows them to be easily applied in practice and we make a suite of tools available for computational analysis with DWTs.

Bioinformatics

Tracking single-cell gene regulation in dynamically controlled environments using an integrated microfluidic and computational setup

Bacteria adapt to changes in their environment by regulating gene expression, often at the level of transcription. However, since the molecular processes underlying gene regulation are subject to thermodynamic and other stochastic fluctuations, gene expression is inherently noisy, and identical cells in a homogeneous environment can display highly heterogeneous expression levels. To study how stochasticity affects gene regulation at the single-cell level, it is crucial to be able to directly follow gene expression dynamics in single cells under changing environmental conditions. Recently developed microfluidic devices, used in combination with quantitative fluorescence time-lapse microscopy, represent a highly promising experimental approach, allowing tracking of lineages of single cells over long time-scales while simultaneously measuring their growth and gene expression. However, current devices do not allow controlled dynamical changes to the environmental conditions which are needed to study gene regulation. In addition, automated analysis of the imaging data from such devices is still highly challenging and no standard software is currently available. To address these challenges, we here present an integrated experimental and computational setup featuring, on the one hand, a new dual-input microfluidic chip which allows mixing and switching between two growth media and, on the other hand, a novel image analysis software which jointly optimizes segmentation and tracking of the cells and allows interactive user-guided fine-tuning of its results. To demonstrate the power of our approach, we study the lac operon regulation in E. coli cells grown in an environment that switches between glucose and lactose, and quantify stochastic lag times and memory at the single cell level.

Systems Biology

Inferring intrinsic and extrinsic noise from a dual fluorescent reporter

Dual fluorescent reporter constructs, which measure gene expression from two identical promoters within the same cell, allow total gene expression noise to be decomposed into an extrinsic component, roughly associated with cell-to-cell fluctuations in cellular component concentrations, and intrinsic noise, roughly associated with inherent stochasticity of the biochemical reactions involved in gene expression [1]. A recent paper by Fu and Pachter presented frequentist statistical estimators for intrinsic and extrinsic noise using data from dual reporters [2]. For comparison, I here present results of a Bayesian analysis of this problem. I show that the orthodox estimators suffer from pathologies such as predicting negative values for a manifestly non-negative quantity, i.e. variance, and show that the Bayesian estimators do not suffer from such pathologies. In addition, I show that the Bayesian analysis automatically identifies that optimal estimates of intrinsic and extrinsic noise depend on a subtle combination of two statistics of the data, allowing for accuracies that are up to twice the accuracy of the orthodox estimators in some parameter regimes.\n\nI hope up this little worked out example contrasting orthodox statistical analysis based on ad hoc estimators with estimators resulting from a Bayesian analysis, will be educational for others in the field. I distribute a Mathematica Notebook with this paper that allows users to easily reproduce all results and figures of the paper.

Bioinformatics

Crunch: Completely Automated Analysis of ChIP-seq Data

Although it has become routine for experimental groups to apply ChIP-seq technology to quantitatively characterize the genome-wide binding of transcription factors (TFs), computational analysis procedures remain far from standardized, making it difficult to meaningfully compare ChIP-seq results across experiments. In addition, while genome-wide binding patterns must ultimately be determined by local constellations of binding sites in the DNA, current analysis is typically limited to a standard search for enriched motifs in ChIP-seq peaks.\n\nHere we present Crunch, a completely automated computational method that performs all ChIP-seq analysis from quality control through read mapping and peak detecting, and integrates comprehensive modeling of the ChIP signal in terms of known and novel binding motifs, quantifying the contribution of each motif, and annotating which combinations of motifs explain each binding peak.\n\nApplying Crunch to 128 ChIP-seq datasets from the ENCODE project we find that TFs naturally separate into solitary TFs, for which a single motif explains the ChIP-peaks, and co-binding TFs for which multiple motifs co-occur within peaks. Moreover, for most datasets the motifs that Crunch identified de novo outperform known motifs and both the set of co-binding motifs and the top motif of solitary TFs are consistent across experiments and cell lines. Crunch is implemented as a web server (crunch.unibas.ch), enabling standardized analysis of any collection of ChIP-seq datasets by simply uploading raw sequencing data. Results are provided both in a graphical interface and as downloadable files.

Bioinformatics

Exploiting variability of single cells to uncover the in vivo hierarchy of miRNA targets

MiRNAs are post-transcriptional repressors of gene expression that may additionally reduce the cell-to-cell variability in protein expression, induce correlations between target expression levels and provide a layer through which targets can influence each others expression as competing RNAs (ceRNAs). Here we combined single cell sequencing of human embryonic kidney cells in which the expression of two distinct miRNAs was induced over a wide range, with mathematical modeling, to estimate Michaelis-Menten (KM)-type constants for hundreds of evolutionarily conserved miRNA targets. These parameters, which we inferred here for the first time in the context of the entire network of endogenous miRNA targets, vary over ~2 orders of magnitude. They reveal an in vivo hierarchy of miRNA targets, defined by the concentration of miRNA-Argonaute complexes at which the targets are most sensitively down-regulated. The data further reveals miRNA-induced correlations in target expression at the single cell level, as well as the response of target noise to the miRNA concentration. The approach is generalizable to other miRNAs and post-transcriptional regulators and provides a deeper understanding of gene expression dynamics.

Systems Biology

Ewing sarcoma breakpoint region 1 prevents transcription-associated genome instability

Ewing Sarcoma break point region 1 (EWSR1) is a multi-functional RNA-binding protein that is involved in many cellular processes, from gene expression to RNA processing and transport. Translocations into its locus lead to chimeric proteins with tumorigenic activity, that lack the RNA binding domain. With crosslinking and immunoprecipitation we have found that EWSR1 binds to intronic regions that are present in polyadenylated nuclear RNAs, which include the translocation-prone region of its own locus. Reduced EWSR1 expression leads to gene expression changes that indicate reduced proliferation. By fluorescence in situ hybridization (FISH) with break-apart probes that flanked the translocation-prone region within the EWSR1 locus we found that reduced EWSR1 expression increases the frequency of split signals, indicative of DNA double strand breaks (DSB). The response in phosphorylated histone H2AX and p53-binding protein 1 (53BP1) double-stained foci to the topoisomerase poison camptothecin in cells treated with a control shRNA and with sh-EWSR1 further suggests that EWSR1 functions in the prevention of DNA DSBs. Our data reveal a new function of the EWSR1 member of the FET family and suggest a connection between the RNA-binding activity of EWSR1 and the instability of its own locus that may play a role in malignancy-associated translocations.

Systems Biology