Search bioRxivSearch

Biology subjects

Lawler, A. J.

Publications and source records attributed to Lawler, A. J..

4 recordsLinked to original sources

Machine learning sequence prioritization for cell type-specific enhancer design

Recent discoveries of extreme cellular diversity in the brain warrant rapid development of technologies to access specific cell populations, enabling characterization of their roles in behavior and in disease states. Available approaches for engineering targeted technologies for new neuron subtypes are low-yield, involving intensive transgenic strain or virus screening. Here, we introduce SNAIL (Specific Nuclear-Anchored Independent Labeling), a new virus-based strategy for cell labeling and nuclear isolation from heterogeneous tissue. SNAIL works by leveraging machine learning and other computational approaches to identify DNA sequence features that confer cell type-specific gene activation and using them to make a probe that drives an affinity purification-compatible reporter gene. As a proof of concept, we designed and validated two novel SNAIL probes that target parvalbumin-expressing (PV) neurons. Furthermore, we show that nuclear isolation using SNAIL in wild type mice is sufficient to capture characteristic open chromatin features of PV neurons in the cortex, striatum, and external globus pallidus. Expansion of this technology has broad applications in cell type-specific observation, manipulation, and therapeutics across species and disease models.

genomics

Predicting lineage-specific differences in open chromatin across dozens of mammalian genomes

BackgroundEvolutionary conservation is an invaluable tool for inferring functional significance in the genome, including regions that are crucial across many species and those that have undergone convergent evolution. Computational methods to test for sequence conservation are dominated by algorithms that examine the ability of one or more nucleotides to align across large evolutionary distances. While these nucleotide alignment-based approaches have proven powerful for protein-coding genes and some non-coding elements, they fail to capture conservation at many enhancers, distal regulatory elements that control spatio-temporal patterns of gene expression. The function of enhancers is governed by a complex, often tissue- and cell type-specific, code that links combinations of transcription factor binding sites and other regulation-related sequence patterns to regulatory activity. Thus, function of orthologous enhancer regions can be conserved across large evolutionary distances, even when nucleotide turnover is high. ResultsWe present a new machine learning-based approach for evaluating enhancer conservation that leverages the combinatorial sequence code of enhancer activity rather than relying on the alignment of individual nucleotides. We first train a convolutional neural network model that is able to predict tissue-specific open chromatin, a proxy for enhancer activity, across mammals. Then, we apply that model to distinguish instances where the genome sequence would predict conserved function versus a loss regulatory activity in that tissue. We present criteria for systematically evaluating model performance for this task and use them to demonstrate that our models accurately predict tissue-specific conservation and divergence in open chromatin between primate and rodent species, vastly out-performing leading nucleotide alignment-based approaches. We then apply our models to predict open chromatin at orthologs of brain and liver open chromatin regions across hundreds of mammals and find that brain enhancers associated with neuron activity and liver enhancers associated with liver regeneration have a stronger tendency than the general population to have predicted lineage-specific open chromatin. ConclusionThe framework presented here provides a mechanism to annotate tissue-specific regulatory function across hundreds of genomes and to study enhancer evolution using predicted regulatory differences rather than nucleotide-level conservation measurements.

genetics

The Regulatory Evolution of the Primate Fine-Motor System

In mammals, fine motor control is essential for skilled behavior, and is subserved by specialized subdivisions of the primary motor cortex (M1) and other components of the brains motor circuitry. We profiled the epigenomic state of several components of the Rhesus macaque motor system, including subdivisions of M1 corresponding to hand and orofacial control. We compared this to open chromatin data from M1 in rat, mouse, and human. We found broad similarities as well as unique specializations in open chromatin regions (OCRs) between M1 subdivisions and other brain regions, as well as species- and lineage-specific differences reflecting their evolutionary histories. By distinguishing shared mammalian M1 OCRs from primate- and human-specific specializations, we highlight gene regulatory programs that could subserve the evolution of skilled motor behaviors such as speech and tool use.

neuroscience

Addiction-associated genetic variants implicate brain cell type- and region-specific cis-regulatory elements in addiction neurobiology

Recent large genome-wide association studies (GWAS) have identified multiple confident risk loci linked to addiction-associated behavioral traits. Genetic variants linked to addiction-associated traits lie largely in non-coding regions of the genome, likely disrupting cis-regulatory element (CRE) function. CREs tend to be highly cell type-specific and may contribute to the functional development of the neural circuits underlying addiction. Yet, a systematic approach for predicting the impact of risk variants on the CREs of specific cell populations is lacking. To dissect the cell types and brain regions underlying addiction-associated traits, we applied LD score regression to compare GWAS to genomic regions collected from human and mouse assays for open chromatin, which is associated with CRE activity. We found enrichment of addiction-associated variants in putative CREs marked by open chromatin in neuronal (NeuN+) nuclei collected from multiple prefrontal cortical areas and striatal regions known to play major roles in reward and addiction. To further dissect the cell type-specific basis of addiction-associated traits, we also identified enrichments in human orthologs of open chromatin regions of mouse neuronal subtypes: cortical excitatory, D1, D2, and PV. Lastly, we developed machine learning models from mouse cell type-specific regions of open chromatin to further dissect human NeuN+ open chromatin regions into cortical excitatory or striatal D1 and D2 neurons and predict the functional impact of addiction-associated genetic variants. Our results suggest that different neuronal subtypes within the reward system play distinct roles in the variety of traits that contribute to addiction. Significance StatementWe combine statistical genetic and machine learning techniques to find that the predisposition to for nicotine, alcohol, and cannabis use behaviors can be partially explained by genetic variants in conserved regulatory elements within specific brain regions and neuronal subtypes of the reward system. This computational framework can flexibly integrate open chromatin data across species to screen for putative causal variants in a cell type-and tissue-specific manner across numerous complex traits.

genomics