Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,693 records · Page 94Linked to original sources

Kinetics of Xist-induced gene silencing can be predicted from combinations of epigenetic and genomic features

To initiate X-chromosome inactivation (XCI), the long non-coding RNA Xist mediates chromosome-wide gene silencing of one X chromosome in female mammals to equalize gene dosage between the sexes. The efficiency of gene silencing, however is highly variable across genes, with some genes even escaping XCI in somatic cells. A genes susceptibility to Xist-mediated silencing appears to be determined by a complex interplay of epigenetic and genomic features; however, the underlying rules remain poorly understood. We have quantified chromosome-wide gene silencing kinetics at the level of the nascent transcriptome using allele-specific Precision nuclear Run-On sequencing (PRO-seq). We have developed a Random Forest machine learning model that can predict the measured silencing dynamics based on a large set of epigenetic and genomic features and tested its predictive power experimentally. While the genomic distance to the Xist locus is the prime determinant of the speed of gene silencing, we find that also pre-marking of gene promoters with polycomb complexes is associated with fast silencing. Moreover, a series of features associated with active transcription and the O-GlcNAc transferase Ogt are enriched at rapidly silenced genes. Our machine learning approach can thus uncover the complex combinatorial rules underlying gene silencing during X inactivation.

genomics

Genomic analysis of natural intra-specific hybrids among Ethiopian isolates of Leishmania donovani

Parasites of the genus Leishmania (Kinetoplastida: Trypanosomatidae) cause widespread and devastating human diseases, ranging from self-healing but disfiguring cutaneous lesions to destructive mucocutaneous presentations or usually fatal visceral disease. Visceral leishmaniasis due to Leishmania donovani is endemic in Ethiopia where it has also been responsible for major epidemics. The presence of hybrid genotypes has been widely reported in surveys of natural populations, genetic variation reported in a number of Leishmania species, and the extant capacity for genetic exchange demonstrated in laboratory experiments. However, patterns of recombination and evolutionary history of admixture that produced these hybrid populations remain unclear, as most of the relevant literature examines only a limited number (typically fewer than 10) genetic loci. Here, we use whole-genome sequence data to investigate Ethiopian L. donovani isolates previously characterised as hybrids by microsatellite and multi-locus sequencing. To date there is only one previous study on a natural population of Leishmania hybrids, based on whole-genome sequence. The current findings demonstrate important differences. We propose hybrids originate from recombination between two different lineages of Ethiopian L. donovani occurring in the same region. Patterns of inheritance are more complex than previously reported with multiple, apparently independent, origins from similar parents that include backcrossing with parental types. Analysis indicates that hybrids are representative of at least three different histories. Furthermore, isolates were highly polysomic at the level of chromosomes with startling differences between parasites recovered from a recrudescent infection from a previously treated individual. The results demonstrate that recombination is a significant feature of natural populations and contributes to the growing body of evidence describing how recombination, and gene flow, shape natural populations of Leishmania.\n\nAuthor SummaryLeishmaniasis is a spectrum of diseases caused by the protozoan parasite Leishmania. It is transmitted by sandfly insect vectors and is responsible for an enormous burden of human suffering. In this manuscript we examine Leishmania isolates from Ethiopia that cause the most serious form of the disease, namely visceral leishmaniasis, which is usually fatal without treatment. Historically the general view was that such parasites reproduce clonally, so that their progeny are genetically identical to the founding cells. This view has changed over time and it is increasingly clear that recombination between genetically different Leishmania parasites occurs. The implication is that new biological traits such as virulence, resistance to drug treatments or the ability to infect new species of sandfly could emerge. The frequency and underlying mechanism of such recombination in natural isolates is poorly understood. Here we perform a detailed whole genome analysis on a cohort of hybrid isolates from Ethiopia together with their potential parents to assess the genetic nature of hybrids in more detail. Results reveal a complex pattern of mating and inbreeding indicative of multiple mating events that has likely shaped the epidemiology of the disease agent. We also show that some hybrids have very different relative amounts of DNA (polysomy) the implications of which are discussed. Together the results contribute to a fuller understanding of the nature of genetic recombination in natural populations of Leishmania.

genomics

The genome of cowpea (Vigna unguiculata Walp.)

Cowpea (Vigna unguiculata [L.] Walp.) is a major crop for worldwide food and nutritional security, especially in sub-Saharan Africa, that is resilient to hot and drought-prone environments. A high-quality assembly of the single-haplotype inbred genome of cowpea IT97K-499-35 was developed by exploiting the synergies between single molecule real-time sequencing, optical and genetic mapping, and a novel assembly reconciliation algorithm. A total of 519 Mb is included in the assembled sequences. Nearly half of the assembled sequence is composed of repetitive elements, which are enriched within recombination-poor pericentromeric regions. A comparative analysis of these elements suggests that genome size differences between Vigna species are mainly attributable to changes in the amount of Gypsy retrotransposons. Conversely, genes are more abundant in more distal, high-recombination regions of the chromosomes; there appears to be more duplication of genes within the NBS-LRR and the SAUR-like auxin superfamilies compared to other warm-season legumes that have been sequenced. A surprising outcome of this study is the identification of a chromosomal inversion of 4.2 Mb among landraces and cultivars, which includes a gene that has been associated in other plants with interactions with the parasitic weed Striga gesnerioides. The genome sequence also facilitated the identification of a putative syntelog for multiple organ gigantism in legumes. A new numbering system has been adopted for cowpea chromosomes based on synteny with common bean (Phaseolus vulgaris).

genomics

High-quality genome assembly and high-density genetic map of asparagus bean

Asparagus bean (Vigna. unguiculata ssp. sesquipedialis), known for its very long and tender green pods, is an important vegetable crop broadly grown in the developing countries. Despite its agricultural and economic values, asparagus bean does not have a high-quality genome assembly for breeding novel agronomic traits. In this study, we reported a high-quality 632.8 Mb assembly of asparagus bean based on the whole genome shotgun sequencing strategy. We also generated a high-density linkage map for asparagus bean, which helped anchor 94.42% of the scaffolds into 11 pseudo-chromosomes. A total of 42,609 protein-coding genes and 3,579 non-protein-coding genes were predicted from the assembly. Taken together, these genomic resources of asparagus bean will facilitate the investigation of economically valuable traits in a variety of legume species, so that the cultivation of these plants would help combat the protein and energy malnutrition in the developing world.

genomics

Is it time to change the reference genome?

The use of the human reference genome has shaped methods and data across modern genomics. This has offered many benefits while creating a few constraints. In the following piece, we outline the history, properties, and pitfalls of the current human reference genome. In a few illustrative analyses, we focus on its use for variant-calling, highlighting its nearness to a "type specimen". We suggest that switching to a consensus reference offers important advantages over the current reference with few disadvantages.

genomics

Interplay of pericentromeric genome organization and chromatin landscape regulates the expression of Drosophila melanogaster heterochromatic genes

Transcription of heterochromatic genes residing within the constitutive heterochromatin is paradoxical to the tenets of the epigenetic code. Drosophila melanogaster heterochromatic genes serve as an excellent model system to understand the mechanisms of their transcriptional regulation. Recent developments in chromatin conformation techniques have revealed that genome organization regulates the transcriptional outputs. Thus, using 5C-seq in S2 cells, we present a detailed characterization of the hierarchical genome organization of Drosophila pericentromeric heterochromatin and its contribution to heterochromatic gene expression. We show that pericentromeric TAD borders are enriched in nuclear Matrix attachment regions while the intra-TAD interactions are mediated by various insulator binding proteins. Heterochromatic genes of similar expression levels cluster into Het TADs which indicates their transcriptional co-regulation. To elucidate how heterochromatic factors, influence the expression of heterochromatic genes, we performed 5C-seq in the HP1a or Su(var)3-9 depleted cells. HP1a or Su(var)3-9 RNAi results in perturbation of global pericentromeric TAD organization but the expression of the heterochromatic genes is minimally affected. Subset of active heterochromatic genes have been shown to have combination of HP1a/H3K9me3 with H3K36me3 at their exons. Interestingly, the knock-down of dMES-4 (H3K36 methyltransferase), downregulates expression of the heterochromatic genes. This indicates that the local chromatin interactions and the combination of heterochromatic factors (HP1a or H3K9me3) along with the H3K36me3 is crucial to drive the expression of heterochromatic genes. Furthermore, dADD1, present near the TSS of the active heterochromatic genes, can bind to both H3K9me3 or HP1a and facilitate the heterochromatic gene expression by regulating the H3K36me3 levels. Therefore, our findings provide mechanistic insights into the interplay of genome organization and chromatin factors at the pericentromeric heterochromatin that regulates Drosophila melanogaster heterochromatic gene expression.

genomics

Single-cell whole-genome sequencing reveals the functional landscape of somatic mutations in B lymphocytes across the human lifespan

Introductory paragraphThe accumulation of mutations in somatic cells have been implicated as a cause of ageing since the 1950s1,2. Yet, attempts to establish a causal relationship between somatic mutations and ageing have been constrained by the lack of methods to directly identify mutational events in primary human tissues. Here we provide detailed, genome-wide mutation frequencies and spectra of human B lymphocytes from healthy individuals across the entire human lifespan, from newborns to centenarians, using a recently developed, highly accurate single-cell whole-genome sequencing method3. We found that the number of somatic mutations increases from <500 per cell in newborns to >3,000 per cell in centenarians. We discovered mutational hotspot regions, some of which, as expected, located at immunoglobulin genes associated with somatic hypermutation. B cell-specific mutation signatures were observed associated with development, ageing or somatic hypermutation (SHM). The SHM signature strongly correlated with the signature found in human chronic lymphocytic leukemia and malignant B-cell lymphomas4, indicating that even in B cells of healthy individuals the potential cancer-causing events are already present. We also identified multiple mutations in sequence features relevant to cellular function, i.e., transcribed genes and gene regulatory regions. Such mutations increased significantly during ageing, but only at approximately half the rate of the genome average, indicating selection against mutations that impact B cell function. This first full characterization of the landscape of somatic mutations in human B lymphocytes indicates that spontaneous somatic mutations accumulating with age can be deleterious and may contribute to both the increased risk for leukemia and the functional decline of B lymphocytes in the elderly.

genomics

Targeted, High-Resolution RNA Sequencing Of Non-Coding Genomic Regions Associated With Neuropsychiatric Functions

The human brain is one of the last frontiers of biomedical research. Genome-wide association studies (GWAS) have succeeded in identifying thousands of haplotype blocks associated with a range of neuropsychiatric traits, including disorders such as schizophrenia, Alzheimers and Parkinsons disease. However, the majority of single nucleotide polymorphisms (SNPs) that mark these haplotype blocks fall within non-coding regions of the genome, hindering their functional validation. While some of these GWAS loci may contain cis-acting regulatory DNA elements such as enhancers, we hypothesized that many are also transcribed into non-coding RNAs that are missing from publicly available transcriptome annotations. Here, we use targeted RNA capture ( RNA CaptureSeq) in combination with nanopore long-read cDNA sequencing to transcriptionally profile 1,023 haplotype blocks across the genome containing non-coding GWAS SNPs associated with neuropsychiatric traits, using post-mortem human brain tissue from three neurologically healthy donors. We find that the majority (62%) of targeted haplotype blocks, including 13% of intergenic blocks, are transcribed into novel, multi-exonic RNAs, most of which are not yet recorded in GENCODE annotations. We validated our findings with short-read RNA-seq, providing orthogonal confirmation of novel splice junctions and enabling a quantitative assessment of the long-read assemblies. Many novel transcripts are supported by independent evidence of transcription including cap analysis of gene expression (CAGE) data and epigenetic marks, and some show signs of potential functional roles. We present these transcriptomes as a preliminary atlas of non-coding transcription in human brain that can be used to connect neurological phenotypes with gene expression.

genomics

Complete genome sequence of Bacillus velezensis JT3-1, a microbial germicide isolated from yak feces

Bacillus velezensis JT3-1 is a probiotic strain isolated from feces of the domestic yak (Bos grunniens) in the Gansu province of China. It has strong antagonistic activity against Listeria monocytogenes, Staphylococcus aureus, Escherichia coli, Salmonella Typhimurium, Mannheimia haemolytica, Staphylococcus hominis, Clostridium perfringens, and Mycoplasma bovis. These properties have made the JT3-1 strain the focus of commercial interest. In this study, we describe the complete genome sequence of JT3-1, with a genome size of 3,929,799 bp, 3761 encoded genes and an average GC content of 46.50%. Whole genome sequencing of Bacillus velezensis JT3-1 will lay a good foundation for elucidation of the mechanisms of its antimicrobial activity, and for its future application.

genomics

Complete assembly of Escherichia coli ST131 genomes using long reads demonstrates antibiotic resistance gene variation within diverse plasmid and chromosomal contexts

The incidence of infections caused by extraintestinal Escherichia coli (ExPEC) is rising globally, which is a major public health concern. ExPEC strains that are resistant to antimicrobials have been associated with excess mortality, prolonged hospital stays and higher healthcare costs. E. coli ST131 is a major ExPEC clonal group worldwide with variable plasmid composition, and has an array of genes enabling antimicrobial resistance (AMR). ST131 isolates frequently encode the AMR genes blaCTX-M-14/15/27, which are often rearranged, amplified and translocated by mobile genetic elements (MGEs). Short DNA reads do not fully resolve the architecture of repetitive elements on plasmids to allow MGE structures encoding blaCTX-M genes to be fully determined. Here, we performed long read sequencing to decipher the genome structures of six E. coli ST131 isolated from six patients. Most long read assemblies generated entire chromosomes and plasmids as single contigs, contrasting with more fragmented assemblies created with short reads alone. The long read assemblies highlighted diverse accessory genomes with blaCTX-M-15, blaCTX-M-14 and blaCTX-M-27 genes identified in three, one and one isolates, respectively. One sample had no blaCTX-M gene. Two samples had chromosomal blaCTX-M-14 and blaCTX-M-15 genes, and the latter was at three distinct locations, likely transposed by the adjacent MGEs: ISEcp1, IS903B and Tn2. This study showed that AMR genes exist in multiple different chromosomal and plasmid contexts even between closely-related isolates within a clonal group such as E. coli ST131. ImportanceDrug-resistant bacteria are a major cause of illness worldwide and a specific subtype called Escherichia coli ST131 cause a significant amount of these infections. ST131 become resistant to treatment by modifying their DNA and by transferring genes among one another via large packages of genes called plasmids, like a game of pass-the-parcel. Tackling infections more effectively requires a better understanding of what plasmids are being exchanged and their exact contents. To achieve this, we applied new high-resolution DNA sequencing technology to six ST131 samples from infected patients and compared the output to an existing approach. A combination of methods shows that drug-resistance genes on plasmids are highly mobile because they can jump into ST131s chromosomes. We found that the plasmids are very elastic and undergo extensive rearrangements even in closely related samples. This application of DNA sequencing technologies illustrates at a new level the highly dynamic nature of ST131 genomes.

genomics

Complete genome sequence and evolution analysis of Psychrobacter sp. YP14 from Gammaridea Gastrointestinal Microbiota of Yap Trench

Psychrobacter sp. YP14, a moderately psychrophilic bacterium belonging to the class Gammaproteobacteria, was isolated from Gammaridea Gastrointestinal Microbiota of Yap Trench. The strain has one circular chromosome of 2,895,311 bp with a 44.66% GC content, consisting of 2333 protein-coding genes, 53 tRNA genes and 9 rRNA genes. Four plasmids were completely assembled and their sizes were 13,712 bp, 19711 bp, 36270 bp, 8194 bp, respectively. In particular, a putative open reading frame (ORF) for dienelactone hydrolase (DLH) related to degradation of chlorinated aromatic hydrocarbons. To get an better understanding of the evolution of Psychrobacter sp. YP14 in this genus, six Psychrobacter strains (G, PRwf-1, DAB_AL43B, AntiMn-1,P11G5, P2G3), with publicly available complete genome, were selected and comparative genomics analysis were performed among them. The closest phylogenetic relationship was identified between strains G and K5 based on 16s gene and ANI (average nucleotide identity) values. Analysis of the pan-genome structure found that YP14 has fewer COG clusters associated with transposons and prophage which indicates fewer sequence rearrangements compared with PRwf-1. Besides, stress response-related genes of strain YP14 demonstrates that it has less strategies to cope with extreme environment, which is consistent with its intestinal habitat. The difference of metabolism and strategies coped with stress response of YP14 are more conducive to the study of microbial survival and metabolic mechanisms in deep sea environment.

genomics

Jomon genome sheds light on East Asian population history

Anatomical modern humans reached East Asia by >40,000 years ago (kya). However, key questions still remain elusive with regard to the route(s) and the number of wave(s) in the dispersal into East Eurasia. Ancient genomes at the edge of East Eurasia may shed light on the detail picture of peopling to East Eurasia. Here, we analyze the whole-genome sequence of a 2.5 kya individual (IK002) characterized with a typical Jomon culture that started in the Japanese archipelago >16 kya. The phylogenetic analyses support multiple waves of migration, with IK002 forming a lineage basal to the rest of the ancient/present-day East Eurasians examined, likely to represent some of the earliest-wave migrants who went north toward East Asia from Southeast Asia. Furthermore, IK002 has the extra genetic affinity with the indigenous Taiwan aborigines, which may support a coastal route of the Jomon-ancestry migration from Southeast Asia to the Japanese archipelago. This study highlight the power of ancient genomics with the isolated population to provide new insights into complex history in East Eurasia.

genomics

SVCurator: A Crowdsourcing app to visualize evidence of structural variants for the human genome

A high quality benchmark for small variants encompassing 88 to 90% of the reference genome has been developed for seven Genome in a Bottle (GIAB) reference samples. However a reliable benchmark for large indels and structural variants (SVs) is yet to be defined. In this study, we manually curated 1235 SVs which can ultimately be used to evaluate SV callers or train machine learning models. We developed a crowdsourcing app - SVCurator - to help curators manually review large indels and SVs within the human genome, and report their genotype and size accuracy.\n\nSVCurator is a Python Flask-based web platform that displays images from short, long, and linked read sequencing data from the GIAB Ashkenazi Jewish Trio son [NIST RM 8391/HG002], We asked curators to assign labels describing SV type (deletion or insertion), size accuracy, and genotype for 1235 putative insertions and deletions sampled from different size bins between 20 and 892,149 bp. The crowdsourced results were highly concordant with 37 out of the 61 curators having at least 78% concordance with a set of expert curators, where there was 93% concordance amongst expert curators. This produced high confidence labels for 935 events. When compared to the heuristic-based draft benchmark SV callset from GIAB, the SVCurator crowdsourced labels were 94.5% concordant with the benchmark set. We found that curators can successfully evaluate putative SVs when given evidence from multiple sequencing technologies.

genomics

An evaluation of pool-sequencing transcriptome-based exon capture for population genomics of non-model species

AO_SCPLOWBSTRACTC_SCPLOWExon capture coupled to high-throughput sequencing constitutes a cost-effective technical solution for addressing specific questions in evolutionary biology by focusing on expressed regions of the genome preferentially targeted by selection. Transcriptome-based capture, a process that can be used to capture the exons of non-model species, is use in phylogenomics. However, its use in population genomics remains rare due to the high costs of sequencing large numbers of indexed individuals across multiple populations. We evaluated the feasibility of combining transcriptome-based capture and the pooling of tissues from numerous individuals for DNA extraction as a cost-effective, generic and robust approach to estimating the variant allele frequencies of any species at the population level. We designed capture probes for [~]5 Mb of chosen de novo transcripts from the Asian ladybird Harmonia axyridis (5,717 transcripts). We called [~]300,000 bi-allelic SNPs for a pool of 36 non-indexed individuals. Capture efficiency was high, and pool-seq was as effective and accurate as individual-seq for detecting variants and estimating allele frequencies. Finally, we also evaluated an approach for simplifying bioinformatic analyses by mapping genomic reads directly to targeted transcript sequences to obtain coding variants. This approach is effective and does not affect the estimation of SNP allele frequencies, except for a small bias close to some exon ends. We demonstrate that this approach can also be used to predict the intron-exon boundaries of targeted de novo transcripts, making it possible to abolish genotyping biases near exon ends.

genomics

Comparative genomic analysis of three salmonid species identifies functional candidate genes involved in resistance to the intracellular bacteria Piscirickettsia salmonis

Piscirickettsia salmonis is the etiological agent of Salmon Rickettsial Syndrome (SRS), and is responsible for considerable economic losses in salmon aquaculture. The bacteria affect coho salmon (CS) (Oncorhynchus kisutch), Atlantic salmon (AS) (Salmo salar) and rainbow trout (RT) (Oncorhynchus mykiss) in several countries, including: Norway, Canada, Scotland, Ireland and Chile. We used Bayesian genome-wide association (GWAS) analyses to investigate the genetic architecture of resistance to P. salmonis in farmed populations of these species. Resistance to SRS was defined as the number of days to death (DD) and as binary survival (BS). A total of 828 CS, 2,130 RT and 2,601 AS individuals were phenotyped and then genotyped using ddRAD sequencing, 57K SNP Affymetrix(R) Axiom(R) and 50K Affymetrix(R) Axiom(R) SNP panels, respectively. Both trait of SRS resistance in CS and RT, appeared to be under oligogenic control. In AS there was evidence of polygenic control of SRS resistance. To identify candidate genes associated with resistance, we applied a comparative genomics approach in which we systematically explored the complete set of genes adjacent to SNPs which explained more than 1% of the genetic variance of resistance in each salmonid species (533 genes in total). Thus, genes were classified based on the following criteria: i) shared function of their protein domains among species, ii) shared orthology among species, iii) proximity to the SNP explaining the highest proportion of the genetic variance and, iv) presence in more than one genomic region explaining more than 1% of the genetic variance within species. Our results allowed us to identify 120 candidate genes belonging to at least one of the four criteria described above. Of these, 21 of them were part of at least two of the criteria defined above and are suggested to be strong functional candidates influencing P. salmonis resistance. These genes are related to diverse biological processes, such as: kinase activity, GTP hydrolysis, helicase activity, lipid metabolism, cytoskeletal dynamics, inflammation and innate immune response, which seem essential in the host response against P. salmonis infection. These results provide fundamental knowledge on the potential functional genes underpinning resistance against P. salmonis in three salmonid species.

genomics

The Complete Chloroplast Genome of Tetraselmis desikacharyi (Chlorodendrophyceae) and Phylogenetic Analysis

Tetraselmis desikacharyi is a marine alga, known as an important plankton for aquaculture as a feed organism. However, the genomic study on this class is rare. Here, we present a complete Chlorodendrophyceae chloroplast genome of T. desikacharyi, belonging to Chlorodendrophyceae with a full length of 149,934bp, characterized by a very small single-copy (SSC) region without any genes and a large inverted repeat (IR) region. A maximum-likelihood (ML) phylogenetic analysis was performed using three kinds of data comprising 50 protein-coding genes, which placed the Chlorodendrophyceae as a deep-diverging lineage of the core Chlorophyta.\n\nT. desikacharyi, characterized by their intense green colored chloroplast, is important for aquaculture as a feed organism, and for studying plankton growth cycles due to their fast growth rate (Arora et al., 2013; Norris et al., 1980). However, only two chloroplast genomes have been published in ...

genomics

Chromatin-lamin B1 interaction promotes genomic compartmentalization and constrains chromatin dynamics

The eukaryotic genome is folded into higher-order conformation accompanied with constrained dynamics for coordinated genome functions. However, the molecular machinery underlying these hierarchically organized chromatin architecture and dynamics remains poorly understood. Here by combining imaging and Hi-C sequencing, we studied the role of lamin B1 in chromatin architecture and dynamics. We found that lamin B1 depletion leads to chromatin redistribution and decompaction. Consequently, the inter-chromosomal interactions and overlap between chromosome territories are increased. Moreover, Hi-C data revealed that lamin B1 is required for the integrity and segregation of chromatin compartments but not for the topologically associating domains (TADs). We further proved that depletion of lamin B1 leads to increased chromatin dynamics, owing to chromatin decompaction and redistribution toward nuclear interior. Taken together, our data suggest that chromatin-lamin B1 interactions promote chromosomal territory segregation and genomic compartmentalization, and confine chromatin dynamics, supporting its crucial role in chromatin higher-order structure and dynamics.

genomics

Integrating regulatory DNA sequence and gene expression to predict genome-wide chromatin accessibility across cellular contexts

MotivationGenome-wide profiles of chromatin accessibility and gene expression in diverse cellular contexts are critical to decipher the dynamics of transcriptional regulation. Recently, convolutional neural networks (CNNs) have been used to learn predictive cis-regulatory DNA sequence models of context-specific chromatin accessibility landscapes. However, these context-specific regulatory sequence models cannot generalize predictions across cell types.\n\nResultsWe introduce multi-modal, residual neural network architectures that integrate cis-regulatory sequence and context-specific expression of trans-regulators to predict genome-wide chromatin accessibility profiles across cellular contexts. We show that the average accessibility of a genomic region across training contexts can be a surprisingly powerful predictor. We leverage this feature and employ novel strategies for training models to enhance genome-wide prediction of shared and context-specific chromatin accessible sites across cell types. We interpret the models to reveal insights into cis and trans regulation of chromatin dynamics across 123 diverse cellular contexts.\n\nAvailabilityThe code is available at https://github.com/kundajelab/ChromDragoNN\n\nContactakundaje@stanford.edu

genomics