Search bioRxiv⌕ Search

Biology subjects

Laverty, K. U.

Publications and source records attributed to Laverty, K. U..

5 recordsLinked to original sources

GHT-SELEX demonstrates unexpectedly high intrinsic sequence specificity and complex DNA binding of many human transcription factors

Precise identification of transcription factor (TF) binding sites is a long standing challenge in human regulatory genomics: TF binding motifs are short and degenerate, while the genome is large. Motif scans, therefore, often produce excessive binding site predictions. By surveying 179 TFs across 25 families using >1,500 cyclic in vitro selection experiments with fragmented, naked, and unmodified genomic DNA - a method we term GHT-SELEX (Genomic HT-SELEX) - we find that many human TFs possess much higher sequence specificity than anticipated. Moreover, genomic binding regions from GHT-SELEX are often surprisingly similar to those obtained in vivo (i.e., ChIP-seq peaks). Contrary to conventional wisdom, we find that high specificity can also be obtained from motif scans, but performance is highly dependent on the derivation and use of the motifs, including accounting for multiple local matches. We also observe alternative engagement of multiple DNA-binding domains within the same protein: long C2H2 zinc finger proteins often utilize modular DNA recognition, engaging different subsets of their DNA-binding domain (DBD) arrays to recognize multiple types of distinct target sites, frequently evolving via internal duplication and divergence of one or more DBDs. Thus, it is common for TFs to possess sufficient intrinsic specificity to delineate a large fraction of in vivo genomic targets, independently of other cellular factors.

genomics↗

Cross-platform DNA motif discovery and benchmarking to explore binding specificities of poorly studied human transcription factors

A DNA sequence pattern, or "motif", is an essential representation of DNA-binding specificity of a transcription factor (TF). Any particular motif model has potential flaws due to shortcomings of the underlying experimental data and computational motif discovery algorithm. As a part of the Codebook/GRECO-BIT initiative, here we evaluated at large scale the cross-platform recognition performance of positional weight matrices (PWMs), which remain popular motif models in many practical applications. We applied ten different DNA motif discovery tools to generate PWMs from the "Codebook" data comprised of 4,237 experiments from five different platforms profiling the DNA-binding specificity of 394 human proteins, focusing on understudied transcription factors of different structural families. For many of the proteins, there was no prior knowledge of a genuine motif. By benchmarking-supported human curation, we constructed an approved subset of experiments comprising about 30% of all experiments and 50% of tested TFs which displayed consistent motifs across platforms and replicates. We present the Codebook Motif Explorer (https://mex.autosome.org), a detailed online catalog of DNA motifs, including the top-ranked PWMs, and the underlying source and benchmarking data. We demonstrate that in the case of high-quality experimental data, most of the popular motif discovery tools detect valid motifs and generate PWMs, which perform well both on genomic and synthetic data. Yet, for each of the algorithms, there were problematic combinations of proteins and platforms, and the basic motif properties such as nucleotide composition and information content offered little help in detecting such pitfalls. By combining multiple PMWs in decision trees, we demonstrate how our setup can be readily adapted to train and test binding specificity models more complex than PWMs. Overall, our study provides a rich motif catalog as a solid baseline for advanced models and highlights the power of the multi-platform multi-tool approach for reliable mapping of DNA binding specificities. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=141 SRC="FIGDIR/small/619379v2_ufig1.gif" ALT="Figure 1"> View larger version (61K): org.highwire.dtl.DTLVardef@79561forg.highwire.dtl.DTLVardef@54c0aorg.highwire.dtl.DTLVardef@1c33f34org.highwire.dtl.DTLVardef@16a93ba_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOGraphical AbstractC_FLOATNO C_FIG

bioinformatics↗

Perspectives on Codebook: sequence specificity of uncharacterized human transcription factors

Gene expression is regulated by transcription factors (TFs), which recognize specific DNA sequence motifs. Several hundred putative human TFs, identified mainly by an apparent DNA-binding domain, lack known binding motifs1, and even for well-characterized TFs, it remains controversial to what degree motifs accurately reflect binding sites in living cells2,3. Here, we describe a systematic effort ("Codebook") to determine the sequence specificity of 332 putative and poorly characterized human TFs. Over 4,000 independent experiments, encompassing multiple in vitro and in vivo assays, produced motifs for just over half (177, or 53%), of which most are unique to a single protein, thereby extending the vocabulary of sequence recognition encoded by human TFs by [~]100 distinct motifs. Moreover, binding motifs identified in vitro are strongly enriched within cellular binding sites. Collectively, the data reveal tens of thousands of previously unknown, conserved, and direct TF binding sites across the human genome. These sites are concentrated in promoter regions, and are predictive of gene expression, illustrating that this new data atlas provides an important step forward in decoding the human genome.

genomics↗

Extensive binding of uncharacterized human transcription factors to genomic dark matter

The functional impact of a large portion of the human genome known as "dark matter DNA", which is composed mainly of repeat sequences, remains enigmatic. The genome also encodes hundreds of putative and poorly characterized transcription factors (TFs). Here, we determined genomic binding locations of 166 poorly characterized human TFs in living cells. Nearly half of them associate strongly with known regulatory regions such as promoters and enhancers, frequently co-localizing with each other at conserved motif matches. The other half often associate with genomic dark matter, however, at largely non-overlapping (i.e., unique) sites, via intrinsic sequence recognition. Fifty-four of the latter half, which we term "Dark TFs", mainly bind within regions of closed chromatin, with each recognizing a unique set of repeat sequences. The Dark TFs include many KZNFs, which are known to bind and silence TEs, and other TFs with apparent repressive functions. By contrast, some may be pioneers: we find that induction of TPRX1, a known regulator of zygotic preimplantation, leads to chromatin opening at many of its binding sites in the dark matter genome. Altogether, our results shed light on a large fraction of poorly characterized human TFs and simultaneously illuminate the diversity of function within the dark matter genome.

genomics↗

Reconstructing the sequence specificities of RNA-binding proteins across eukaryotes

RNA-binding proteins (RBPs) are key regulators of gene expression. Here, we introduce EuPRI (Eukaryotic Protein-RNA Interactions) - a freely available resource of RNA motifs for 34,736 RBPs from 690 eukaryotes. EuPRI includes in vitro binding data for 504 RBPs, including newly collected RNAcompete data for 174 RBPs, along with thousands of reconstructed motifs. We reconstruct these motifs with a new computational platform -- Joint Protein-Ligand Embedding (JPLE) -- which can detect distant homology relationships and map specificity-determining peptides. EuPRI quadruples the number of known RBP motifs, expanding the motif repertoire across all major eukaryotic clades, and assigning motifs to the majority of human RBPs. EuPRI drastically improves knowledge of RBP motifs in flowering plants. For example, it increases the number of Arabidopsis thaliana RBP motifs 7-fold, from 14 to 105. EuPRI also has broad utility for inferring post-transcriptional function and evolutionary relationships. We demonstrate this by predicting a role for 12 Arabidopsis thaliana RBPs in RNA stability and identifying rapid and recent evolution of post-transcriptional regulatory networks in worms and plants. In contrast, the vertebrate RNA motif set has remained relatively stable after its drastic expansion between the metazoan and vertebrate ancestors. EuPRI represents a powerful resource for the study of gene regulation across eukaryotes.

genomics↗