Search bioRxiv⌕ Search

Biology subjects

Kribelbauer-Swietek, J. F.

Publications and source records attributed to Kribelbauer-Swietek, J. F..

3 recordsLinked to original sources

EXTRA-seq: a genome-integrated extended massively parallel reporter assay to quantify enhancer-promoter communication

Precise control of gene expression is essential for cellular function, but the mechanisms by which enhancers communicate with promoters to coordinate this process are not fully understood. While sequence-based deep learning models show promise in predicting enhancer-driven gene expression, experimental validation and human-interpretable mechanistic insights lag behind. Here, we present EXTRA-seq, a novel EXTended Reporter Assay followed by sequencing designed to quantify enhancer activity in endogenous contexts over kilobase-scale distances. We demonstrate that EXTRA-seq can be targeted to disease-relevant loci and captures expression changes at the resolution of individual transcription factor binding sites, enabling mechanistic discovery. Using engineered synthetic enhancer-promoter combinations, we reveal that the TATA-box acts as a dynamic range amplifier, modulating expression levels in function of enhancer strength. Importantly, we find that integrating state-of-the-art deep learning models with plasmid-based enhancer assays improves the prediction of gene expression as measured by EXTRA-seq. These findings open new avenues for predictive modeling and therapeutic applications. Overall, our work provides a powerful experimental platform to interrogate the complex interplay between enhancers and promoters, bridging the gap between in silico predictions and human-interpretable biological mechanisms.

genomics↗

Cross-platform DNA motif discovery and benchmarking to explore binding specificities of poorly studied human transcription factors

A DNA sequence pattern, or "motif", is an essential representation of DNA-binding specificity of a transcription factor (TF). Any particular motif model has potential flaws due to shortcomings of the underlying experimental data and computational motif discovery algorithm. As a part of the Codebook/GRECO-BIT initiative, here we evaluated at large scale the cross-platform recognition performance of positional weight matrices (PWMs), which remain popular motif models in many practical applications. We applied ten different DNA motif discovery tools to generate PWMs from the "Codebook" data comprised of 4,237 experiments from five different platforms profiling the DNA-binding specificity of 394 human proteins, focusing on understudied transcription factors of different structural families. For many of the proteins, there was no prior knowledge of a genuine motif. By benchmarking-supported human curation, we constructed an approved subset of experiments comprising about 30% of all experiments and 50% of tested TFs which displayed consistent motifs across platforms and replicates. We present the Codebook Motif Explorer (https://mex.autosome.org), a detailed online catalog of DNA motifs, including the top-ranked PWMs, and the underlying source and benchmarking data. We demonstrate that in the case of high-quality experimental data, most of the popular motif discovery tools detect valid motifs and generate PWMs, which perform well both on genomic and synthetic data. Yet, for each of the algorithms, there were problematic combinations of proteins and platforms, and the basic motif properties such as nucleotide composition and information content offered little help in detecting such pitfalls. By combining multiple PMWs in decision trees, we demonstrate how our setup can be readily adapted to train and test binding specificity models more complex than PWMs. Overall, our study provides a rich motif catalog as a solid baseline for advanced models and highlights the power of the multi-platform multi-tool approach for reliable mapping of DNA binding specificities. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=141 SRC="FIGDIR/small/619379v2_ufig1.gif" ALT="Figure 1"> View larger version (61K): org.highwire.dtl.DTLVardef@79561forg.highwire.dtl.DTLVardef@54c0aorg.highwire.dtl.DTLVardef@1c33f34org.highwire.dtl.DTLVardef@16a93ba_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOGraphical AbstractC_FLOATNO C_FIG

bioinformatics↗

Identification of methylation-sensitive human transcription factors using meSMiLE-seq

Transcription factors (TFs) are key players in eukaryotic gene regulation, but the DNA binding specificity of many TFs remains unknown. Here, we assayed 284 mostly poorly characterized, putative human TFs using selective microfluidics-based ligand enrichment followed by sequencing (SMiLE-seq), revealing 72 new DNA binding motifs. To investigate whether some of the 158 TFs for which we did not find motifs preferably bind epigenetically modified DNA (i.e. methylated CG dinucleotides), we developed methylation-sensitive SMiLE-seq (meSMiLE-seq). This microfluidic assay simultaneously probes the affinity of a protein to methylated and unmethylated DNA, augmenting the capabilities of the original method to infer methylation-aware binding sites. We assayed 114 TFs with meSMiLE-seq and identified DNA-binding models for 48 proteins, including the known methylation-sensitive binding modes for POU5F1 and RFX5. For 11 TFs, binding to methylated DNA was preferred or resulted in the discovery of alternative, methylation-dependent motifs (e.g. PRDM13), while aversion towards methylated sequences was found for 13 TFs (e.g. USF3). Finally, we uncovered a potential role for ZHX2 as a putative binder of Z-DNA, a left-handed helical DNA structure which is adopted more frequently upon CpG methylation. Altogether, our study significantly expands the human TF codebook by identifying DNA binding motifs for 98 TFs, while providing a versatile platform to quantitatively assay the impact of DNA modifications on TF binding.

genomics↗