Search bioRxiv⌕ Search

Biology subjects

Suzuki, P.

Publications and source records attributed to Suzuki, P..

5 recordsLinked to original sources

Prediction and design of transcriptional repressor domains with large-scale mutational scans and deep learning

Regulatory proteins have evolved diverse repressor domains (RDs) to enable precise context-specific repression of transcription. However, our understanding of how sequence variation impacts the functional activity of RDs is limited. To address this gap, we generated a high-throughput mutational scanning dataset measuring the repressor activity of 115,000 variant sequences spanning more than 50 RDs in human cells. We identified thousands of clinical variants with loss or gain of repressor function, including TWIST1 HLH variants associated with Saethre-Chotzen syndrome and MECP2 domain variants associated with Rett syndrome. We also leveraged these data to annotate short linear interacting motifs (SLiMs) that are critical for repression in disordered RDs. Then, we designed a deep learning model called TENet (Transcriptional Effector Network) that integrates sequence, structure and biochemical representations of sequence variants to accurately predict repressor activity. We systematically tested generalization within and across domains with varying homology using the mutational scanning dataset. Finally, we employed TENet within a directed evolution sequence editing framework to tune the activity of both structured and disordered RDs and experimentally test thousands of designs. Our work highlights critical considerations for future dataset design and model training strategies to improve functional variant prioritization and precision design of synthetic regulatory proteins.

genomics↗

High-throughput affinity measurements of direct interactions between activation domains and co-activators

Sequence-specific activation by transcription factors is essential for gene regulation1,2. Key to this are activation domains, which often fall within disordered regions of transcription factors3,4 and recruit co-activators to initiate transcription5. These interactions are difficult to characterize via most experimental techniques because they are typically weak and transient6,7. Consequently, we know very little about whether these interactions are promiscuous or specific, the mechanisms of binding, and how these interactions tune the strength of gene activation. To address these questions, we developed a microfluidic platform for expression and purification of hundreds of activation domains in parallel followed by direct measurement of co-activator binding affinities (STAMMPPING, for Simultaneous Trapping of Affinity Measurements via a Microfluidic Protein-Protein INteraction Generator). By applying STAMMPPING to quantify direct interactions between eight co-activators and 204 human activation domains (>1,500 Kds), we provide the first quantitative map of these interactions and reveal 334 novel binding pairs. We find that the metazoan-specific co-activator P300 directly binds >100 activation domains, potentially explaining its widespread recruitment across the genome to influence transcriptional activation. Despite sharing similar molecular properties (e.g. enrichment of negative and hydrophobic residues), activation domains utilize distinct biophysical properties to recruit certain co-activator domains. Co-activator domain affinity and occupancy are well-predicted by analytical models that account for multivalency, and in vitro affinities quantitatively predict activation in cells with an ultrasensitive response. Not only do our results demonstrate the ability to measure affinities between even weak protein-protein interactions in high throughput, but they also provide a necessary resource of over 1,500 activation domain/co-activator affinities which lays the foundation for understanding the molecular basis of transcriptional activation.

biophysics↗

High-throughput thermodynamic and kinetic measurements of transcription factor/DNA mutations reveal how conformational heterogeneity can shape motif selectivity

Transcription factors (TFs) bind DNA sequences with a range of affinities, yet the mechanisms determining energetic differences between high- and low-affinity sequences ( selectivity) remain poorly understood. Here, we investigated two basic helix-loop-helix TFs, MAX (H. sapiens) and Pho4 (S. cerevisiae), that bind the same high-affinity sequence with highly similar nucleotide-contacting residues and bound structures but are differentially selective for non-cognate sequences. By measuring >1700 Kds and >500 rate constants for Pho4 and MAX mutant libraries binding multiple DNA sequences and comparing these measurements with thermodynamic and kinetic models, we identify the biophysical mechanisms by which changes to TF sequence alter both bound and unbound conformational ensembles to shape specificity landscapes. These results highlight the importance of conformational heterogeneity in determining sequence specificity and selectivity and can guide future efforts to engineer nucleic acid-binding proteins with enhanced selectivity.

biophysics↗

Large-scale mapping and systematic mutagenesis of human transcriptional effector domains

Human gene expression is regulated by over two thousand transcription factors and chromatin regulators1,2. Effector domains within these proteins can activate or repress transcription. However, for many of these regulators we do not know what type of transcriptional effector domains they contain, their location in the protein, their activation and repression strengths, and the amino acids that are necessary for their functions. Here, we systematically measure the transcriptional effector activity of >100,000 protein fragments (each 80 amino acids long) tiling across most chromatin regulators and transcription factors in human cells (2,047 proteins). By testing the effect they have when recruited at reporter genes, we annotate 307 new activation domains and 592 new repression domains, a [~]5-fold increase over the number of previously annotated effectors3,4. Complementary rational mutagenesis and deletion scans across all the effector domains reveal aromatic and/or leucine residues interspersed with acidic, proline, serine, and/or glutamine residues are necessary for activation domain activity. Additionally, the majority of repression domain sequences contain either sites for SUMOylation, short interaction motifs for recruiting co-repressors, or are structured binding domains for recruiting other repressive proteins. Surprisingly, we discover bifunctional domains that can both activate and repress and can dynamically split a cell population into high- and low-expression subpopulations. Our systematic annotation and characterization of transcriptional effector domains provides a rich resource for understanding the function of human transcription factors and chromatin regulators, engineering compact tools for controlling gene expression, and refining predictive computational models of effector domain function.

systems biology↗

High-throughput discovery and characterization of human transcriptional effectors

Thousands of proteins localize to the nucleus; however, it remains unclear which contain transcriptional effectors. Here, we develop HT-recruit - a pooled assay where protein libraries are recruited to a reporter, and their transcriptional effects are measured by sequencing. Using this approach, we measure gene silencing and activation for thousands of domains. We find a relationship between repressor function and evolutionary age for the KRAB domains, discover Homeodomain repressor strength is collinear with Hox genetic organization, and identify activities for several Domains of Unknown Function. Deep mutational scanning of the CRISPRi KRAB maps the co-repressor binding surface and identifies substitutions that improve stability/silencing. By tiling 238 proteins, we find repressors as short as 10 amino acids. Finally, we report new activator domains, including a divergent KRAB. Together, these results provide a resource of 600 human proteins containing effectors and demonstrate a scalable strategy for assigning functions to protein domains.

genetics↗