Search bioRxiv⌕ Search

Biology subjects

Smaruj, P. N.

Publications and source records attributed to Smaruj, P. N..

2 recordsLinked to original sources

Sequence design for three-dimensional genome folding using Akita Semifreddo

Mammalian genomes display complex three-dimensional organization which is crucial for downstream processes like gene regulation. Local features of genome organization are largely driven by loop extrusion and manifest as boundaries, dots, and flames in genome contact maps. Still, the rational design of DNA sequences that produce desired folding patterns has not been demonstrated. Here, we present Akita Semifreddo, a framework that enables the rational in silico design of DNA sequences with programmable 3D folding outcomes. This combines a computationally efficient "half-frozen" version of the AkitaV2 genome folding model with the Ledidi sequence optimizer. We systematically demonstrate that this approach spans the full repertoire of known local folding features. We show that [~]2 kb synthetic sequences can be designed to induce boundaries, dots, and flames at desired strengths, with CTCF motif configurations consistent with their known mechanistic bases. We further demonstrate that weak boundaries can be designed through transcription-associated sequence features alone, without introducing CTCF motifs, and that strong boundaries can be suppressed by introducing SINE B2 retroelement-like sequences. Collectively, these results reveal a many-to-one relationship between DNA sequence and folding outcomes and uncover the biological basis of sequence features leveraged by our model. In short, Akita Semifreddo provides a platform for dissecting the sequence grammar of three-dimensional chromatin architecture and engineering synthetic regulatory landscapes.

genomics↗

Interpreting the CTCF-mediated sequence grammar of genome folding with AkitaV2

Interphase mammalian genomes are folded in 3D with complex locus-specific patterns that impact gene regulation. CTCF (CCCTC-binding factor) is a key architectural protein that binds specific DNA sites, halts cohesin-mediated loop extrusion, and enables long-range chromatin interactions. There are hundreds of thousands of annotated CTCF-binding sites in mammalian genomes; disruptions of some result in distinct phenotypes, while others have no visible effect. Despite their importance, the determinants of which CTCF sites are necessary for genome folding and gene regulation remain unclear. Here, we update and utilize Akita, a convolutional neural network model, to extract the sequence preferences and grammar of CTCF contributing to genome folding. Our analyses of individual CTCF sites reveal four predictions: (i) only a small fraction of genomic sites are impactful, (ii) insulation strength is highly dependent on sequences flanking the core CTCF binding motif, (iii) core and flanking sequences are broadly compatible, and (iv) core and flanking nucleotides contribute largely additively to overall strength. Our analysis of collections of CTCF sites make two predictions for multi-motif grammar: (i) insulation strength depends on the number of CTCF sites within a cluster, and (ii) pattern formation is governed by the orientation and spacing of these sites, rather than any inherent specialization of the CTCF motifs themselves. In sum, we present a framework for using neural network models to probe the sequences instructing genome folding and provide a number of predictions to guide future experimental inquiries. Author SummaryMammalian genomes are spatially organized in 3D with profound consequences for all processes involving DNA. CTCF is a key genome organizer, recognizing numerous sites and creating a variety of contact patterns across the genome. Despite the importance of CTCF, the sequence determinants and grammar of how individual sites collectively instruct genome folding remain unclear. This work leverages the ability of Akita, a deep neural network, to make high-throughput predictions for genome folding after DNA sequence perturbations. Using Akita, we make several experimentally testable predictions. First, only a minority of annotated sites individually impact folding, and flanking DNA sequences greatly modulate their impact. Second, multiple sites together influence folding based on their number, orientation, and spacing. In sum, we provide a roadmap for interpreting neural networks to better understand genome folding and important considerations for the design of experiments.

genomics↗