Search bioRxiv⌕ Search

Biology subjects

Haubrock, M.

Publications and source records attributed to Haubrock, M..

2 recordsLinked to original sources

TFClassPredict: A Novel Deep Learning Framework for Transcription Factor Binding Site Analysis Using Evolutionarily Conserved DNA-Binding Domain Annotations

Defining the cis-regulatory code remains one of the central challenges of modern genomics, requiring the reliable association of transcription factor binding sites (TFBSs) with their cognate transcription factors (TFs) from DNA sequence information. While deep learning methods outperform conventional position weight matrix (PWM)-based approaches, both typically formulate TFBS prediction as a bound-versus-unbound decision rather than resolving the competitive binding potential among TFs at a given genomic location. This is further complicated by the widespread sharing of DNA-binding domains (DBDs) among TFs, which produces overlapping binding preferences, rendering TFBS assignment at the individual TF level inherently ambiguous. Consequently, we present TFClassPredict, a DNABERT-based framework that reframes TFBS prediction as a multi-class classification problem across 23 DBD-classes, discriminating each against all remaining DBD-classes, directly enabling the resolution of binding events between DBD-classes. Trained on high-confidence directly bound TFBSs and exploiting DBD-DNA co-evolution, TFClassPredict naturally resolves the ambiguity arising from shared DBD architectures. We demonstrate that TFClassPredict achieves robust DBD-class level classification, outperforming PWM-based approaches and benchmarked deep learning architectures with predictions mapping to biologically defined DBD-classes. Beyond classification, TFClassPredict improves PWM-based pipelines, three-dimensional genome interaction prediction, and allows genome-wide DBD-class annotation, while recapitulating known lineage-specifying TF programs from immune cell ATAC-seq data.

bioinformatics↗

The importance of DNA sequence for nucleosome positioning in the process of transcriptional regulation

Nucleosome positioning is a key factor for transcriptional regulation. Nucleosomes regulate the dynamic accessibility of chromatin and interact with the transcription machinery at every stage. Influences to steer nucleosome positioning are diverse, and the according importance of the DNA sequence in contrast to active chromatin remodeling has been subject of long discussion. In this study, we evaluate the functional role of DNA sequence for all major elements along the process of transcription. We developed a random forest classifier based on local DNA structure that assesses the sequence-intrinsic support for nucleosome positioning. On this basis, we created a simple data resource that we applied genome-wide to the human genome. In our comprehensive analysis, we found a special role of DNA in mediating the competition of nucleosomes with cis-regulatory elements, in enabling steady transcription, for positioning of stable nucleosomes in exons and for repelling nucleosomes during transcription termination. In contrast, we relate these findings to concurrent processes that generate strongly positioned nucleosomes in vivo that are not mediated by sequence, such as energy-dependent remodeling of chromatin. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=150 SRC="FIGDIR/small/550795v3_ufig1.gif" ALT="Figure 1"> View larger version (45K): org.highwire.dtl.DTLVardef@46167corg.highwire.dtl.DTLVardef@16e6180org.highwire.dtl.DTLVardef@1c34758org.highwire.dtl.DTLVardef@180ea7d_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗