Search bioRxiv⌕ Search

Biology subjects

Duncan, A. G.

Publications and source records attributed to Duncan, A. G..

3 recordsLinked to original sources

A Proximal Sox2 Enhancer Cluster is Required for the Anterior Regional Identity of Neural Progenitors

Embryonic development depends on spatially and temporally orchestrated gene regulatory networks. Expressed in neural stem and progenitor cells (NSPCs), the transcription factor sex-determining region Y box 2 (Sox2) is critical for embryogenesis and stem cell maintenance in neural development. Whereas Sox2 is regulated by a distal cluster of enhancers in embryonic stem cells (ESCs), enhancers closer to the gene have been implicated in Sox2 transcriptional regulation in the neural lineage. Using functional genomics data, and deletion analysis we show that a downstream enhancer cluster regulates Sox2 transcription in NSPCs derived from mouse ESCs. By generating allelic mutants using CRISPR-Cas9 mediated deletions, we show that this proximal enhancer cluster, termed Sox2 regulatory regions 2-18 (SRR2-18), is a cis regulator of Sox2 transcription during neural differentiation. Transcriptome analyses demonstrate that loss of even one copy of SRR2-18 disrupts the region-specific identity of NSPCs. Biallelic deletion of this Sox2 neural enhancer cluster causes reduced SOX2 protein, less frequent interaction with transcriptional machinery, and leads to perturbed chromatin accessibility genome-wide further affecting the expression of neurodevelopmental and anterior-posterior regionalization genes. Furthermore, homozygous NSPC deletants exhibit self-renewal defects and impaired differentiation into cell types found in the brain. Altogether, our data define a cis-regulatory enhancer cluster controlling Sox2 transcription in NSPCs and highlight the sensitivity of neural differentiation processes to decreased Sox2 transcription, which influences their differentiation into posterior neural fates, specifically the caudal neural tube.

genetics↗

Improving the performance of supervised deep learning for regulatory genomics using phylogenetic augmentation

Structured abstractO_ST_ABSMotivationC_ST_ABSSupervised deep learning is used to model the complex relationship between genomic sequence and regulatory function. Understanding how these models make predictions can provide biological insight into regulatory functions. Given the complexity of the sequence to regulatory function mapping (the cis-regulatory code), it has been suggested that the genome contains insufficient sequence variation to train models with suitable complexity. Data augmentation is a widely used approach to increase the data variation available for model training, however current data augmentation methods for genomic sequence data are limited. ResultsInspired by the success of comparative genomics, we show that augmenting genomic sequences with evolutionarily related sequences from other species, which we term phylogenetic augmentation, improves the performance of deep learning models trained on regulatory genomic sequences to predict high-throughput functional assay measurements. Additionally, we show that phylogenetic augmentation can rescue model performance when the training set is down-sampled and permits deep learning on a real-world small dataset, demonstrating that this approach improves experimental data efficiency. Overall, this data augmentation method represents a solution for improving model performance that is applicable to many supervised deep learning problems in genomics. Availability and implementationThe open-source GitHub repository agduncan94/phylogenetic_augmentation_paper includes the code for rerunning the analyses here and recreating the figures. Contactalan.moses@utoronto.ca

bioinformatics↗

Functional similarity of non-coding regions is revealed in phylogenetic average motif score representations

Here we frame the cis-regulatory code (that connects the regulatory functions of non-coding regions, such as promoters and UTRs, to their DNA sequences) as a representation building problem. Representation learning has emerged as a new approach to understand function of DNA and proteins, by projecting sequences into high-dimensional feature spaces, where the features are learned from data by a neural network. Inspired by these approaches, we seek to define a feature space where non-coding regions with similar regulatory functions are nearby each other. As a first attempt, we engineered features based on matches to biochemically characterized regulatory motifs in the DNA sequences of non-coding regions. Remarkably, we found that functionally similar promoters and 3 UTRs could be grouped together in a feature space defined by simple averages of the best match scores in (unaligned) orthologous non-coding regions, which we refer to as phylogenetic average motif scores. Perhaps most important, because this feature space is based on known motifs and not fit to any data, it is fully interpretable and not limited to any particular cell type or experimental context. We find that we can read off known regulatory relationships and evolutionary rewiring from visualizations of phylogenetic average motif score representations, and that predicted regulatory interactions based on neighbors in the feature space are borne out in transcription factor deletion experiments. Phylogenetic averages of match scores to known motifs is a baseline for representation learning applied to non-coding sequences, and may continue to improve as databases of motifs become more complete.

genomics↗