Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,729 records · Page 96Linked to original sources

Chromatin-lamin B1 interaction promotes genomic compartmentalization and constrains chromatin dynamics

The eukaryotic genome is folded into higher-order conformation accompanied with constrained dynamics for coordinated genome functions. However, the molecular machinery underlying these hierarchically organized chromatin architecture and dynamics remains poorly understood. Here by combining imaging and Hi-C sequencing, we studied the role of lamin B1 in chromatin architecture and dynamics. We found that lamin B1 depletion leads to chromatin redistribution and decompaction. Consequently, the inter-chromosomal interactions and overlap between chromosome territories are increased. Moreover, Hi-C data revealed that lamin B1 is required for the integrity and segregation of chromatin compartments but not for the topologically associating domains (TADs). We further proved that depletion of lamin B1 leads to increased chromatin dynamics, owing to chromatin decompaction and redistribution toward nuclear interior. Taken together, our data suggest that chromatin-lamin B1 interactions promote chromosomal territory segregation and genomic compartmentalization, and confine chromatin dynamics, supporting its crucial role in chromatin higher-order structure and dynamics.

genomics

Integrating regulatory DNA sequence and gene expression to predict genome-wide chromatin accessibility across cellular contexts

MotivationGenome-wide profiles of chromatin accessibility and gene expression in diverse cellular contexts are critical to decipher the dynamics of transcriptional regulation. Recently, convolutional neural networks (CNNs) have been used to learn predictive cis-regulatory DNA sequence models of context-specific chromatin accessibility landscapes. However, these context-specific regulatory sequence models cannot generalize predictions across cell types.\n\nResultsWe introduce multi-modal, residual neural network architectures that integrate cis-regulatory sequence and context-specific expression of trans-regulators to predict genome-wide chromatin accessibility profiles across cellular contexts. We show that the average accessibility of a genomic region across training contexts can be a surprisingly powerful predictor. We leverage this feature and employ novel strategies for training models to enhance genome-wide prediction of shared and context-specific chromatin accessible sites across cell types. We interpret the models to reveal insights into cis and trans regulation of chromatin dynamics across 123 diverse cellular contexts.\n\nAvailabilityThe code is available at https://github.com/kundajelab/ChromDragoNN\n\nContactakundaje@stanford.edu

genomics

Loss-of-function tolerance of enhancers in the human genome

Previous studies have surveyed the potential impact of loss-of-function (LoF) variants and identified LoF-tolerant protein-coding genes. However, the tolerance of human genomes to losing enhancers has not yet been evaluated. Here we present the catalog of LoF-tolerant enhancers using structural variants from whole-genome sequences. Using a conservative approach, we estimate that each individual human genome possesses at least 28 LoF-tolerant enhancers on average. We assessed the properties of LoF-tolerant enhancers in a unified regulatory network constructed by integrating tissue-specific enhancers and gene-gene interactions. We find that LoF-tolerant enhancers are more tissue-specific and regulate fewer and more dispensable genes. They are enriched in immune-related cells while LoF-intolerant enhancers are enriched in kidney and brain/neuronal stem cells. We developed a supervised learning approach to predict the LoF-tolerance of enhancers, which achieved an AUROC of 96%. We predict 5,677 more enhancers would be likely tolerant to LoF and 75 enhancers that would be highly LoF-intolerant. Our predictions are supported by known set of disease enhancers and novel deletions from PacBio sequencing. The LoF-tolerance scores provided here will serve as an important reference for disease studies.

genomics

The genome of the endangered dryas monkey provides new insights into the evolutionary history of the vervets

Genomic data can be a powerful tool for inferring ecology, behaviour and conservation needs of highly elusive species, particularly when other sources of information are hard to come by. Here we focus on the Dryas monkey (Cercopithecus dryas), an endangered primate endemic to the Congo Basin with cryptic behaviour and possibly less than 250 remaining individuals. Using whole genome sequencing data we show that the Dryas monkey represents a sister lineage to the vervets (Chlorocebus sp.) and has diverged from them around 1.4 million years ago with additional bi-directional gene flow [~]750,000 - [~]500,000 years ago that has likely involved the crossing of the Congo River. Together with evidence of gene flow across the Congo River in bonobos and okapis, our results suggest that the fluvial topology of the Congo River might have been more dynamic than previously recognised. Despite the presence of several homozygous loss-of-function mutations in genes associated with sperm mobility and immunity, we find high genetic diversity and low levels of inbreeding and genetic load in the studied Dryas monkey individual. This suggests that the current population carries sufficient genetic variability for long-term survival and might be larger than currently recognised. We thus provide an example of how genomic data can directly improve our understanding of highly elusive species.

genomics

Methylation content sensitive enzyme ddRAD (MCSeEd): a reference-free, whole genome profiling system to address cytosine/ adenine methylation changes

Methods for investigating DNA methylation nowadays either require a reference genome and high coverage, or investigate only CG methylation. Moreover, no large-scale analysis can be performed for N6-methyladenosine (6mA). Here we describe the methylation content sensitive enzyme double-digest restriction-site-associated DNA (ddRAD) technique (MCSeEd), a reduced-representation, reference-free, cost-effective approach for characterizing whole genome methylation patterns across different methylation contexts (e.g., CG, CHG, CHH, 6mA). MCSeEd can also detect genetic variations among hundreds of samples. MCSeEd is based on parallel restrictions carried out by combinations of methylation insensitive and sensitive endonucleases, followed by next-generation sequencing. Moreover, we present a robust bioinformatic pipeline (available at https://bitbucket.org/capemaster/mcseed/src/master/) for differential methylation analysis combined with single nucleotide polymorphism calling without or with a reference genome.

genomics

Animal, fungi, and plant genome sequences harbour different non-canonical splice sites

Most protein encoding genes in eukaryotes contain introns which are interwoven with exons. After transcription, introns need to be removed in order to generate the final mRNA which can be translated into an amino acid sequence. Precise excision of introns by the spliceosome requires conserved dinucleotides which mark the splice sites. However, there are variations of the highly conserved combination of GT at the 5 end and AG at the 3 end of an intron in the genome. GC-AG and AT-AC are two major non-canonical splice site combinations which have been known for years. During the last years, various minor non-canonical splice site combinations were detected with numerous dinucleotide permutations. Here we expand systematic investigations of non-canonical splice site combinations in plants to all eukaryotes by analysing fungal and animal genome sequences. Comparisons of splice site combinations between these three kingdoms revealed several differences such as a substantially increased CT-AC frequency in fungal genome sequences. Canonical GT-AG splice site combinations in antisense transcripts could be one explanation for this observation. In addition, high numbers of GA-AG splice site combinations were observed in Eurytemora affinis and Oikopleura dioica. A variant in one U1 snRNA isoform might allow the recognition of GA as 5 splice site. In depth investigation of splice site usage based on RNA-Seq read mappings indicates a generally higher flexibility of the 3 splice site compared to the 5 splice site across animals, fungi, and plants.

genomics

Insights about variation in meiosis from 31,228 human sperm genomes

Meiosis, while critical for reproduction, is also highly variable and error prone: crossover rates vary among humans and individual gametes, and chromosome nondisjunction leads to aneuploidy, a leading cause of miscarriage. To study variation in meiotic outcomes within and across individuals, we developed a way to sequence many individual sperm genomes at once. We used this method to sequence the genomes of 31,228 gametes from 20 sperm donors, identifying 813,122 crossovers, 787 aneuploid chromosomes, and unexpected genomic anomalies. Different sperm donors varied four-fold in the frequency of aneuploid sperm, and aneuploid chromosomes gained in meiosis I had 36% fewer crossovers than corresponding non-aneuploid chromosomes. Diverse recombination phenotypes were surprisingly coordinated: donors with high average crossover rates also made a larger fraction of their crossovers in centromere-proximal regions and placed their crossovers closer together. These same relationships were also evident in the variation among individual gametes from the same donor: sperm with more crossovers tended to have made crossovers closer together and in centromere-proximal regions. Variation in the physical compaction of chromosomes could help explain this coordination of meiotic variation across chromosomes, gametes, and individuals.

genomics

Population structure of Apodemus flavicollis and comparison to Apodemus sylvaticus in northern Poland based on whole-genome genotyping with RAD-seq

BackgroundMice of the genus Apodemus are one the most common mammals in the Palaearctic region. Despite their broad range and long history of ecological observations, there are no whole-genome data available for Apodemus, hindering our ability to further exploit the genus in evolutionary and ecological genomics context. ResultsHere we present results from the double-digest restriction site-associated DNA sequencing (ddRAD-seq) on 72 individuals of A. flavicollis and 10 A. sylvaticus from four populations, sampled across 500 km distance in northern Poland. Our data present clear genetic divergence of the two species, with average p-distance, based on 21377 common loci, of 1.51% and a mutation rate of 0.0011 - 0.0019 substitutions per site per million years. We provide a catalogue of 117 highly divergent loci that enable genetic differentiation of the two species in Poland and to a large degree of 20 unrelated samples from several European countries and Tunisia. We also show evidence of admixture between the three A. flavicollis populations but demonstrate that they have negligible average population structure, with largest pairwise FST < 0.086. ConclusionOur study demonstrates the feasibility of ddRAD-seq in Apodemus and provides the first insights into the population genomics of the species.

genomics

Comprehensive genomic analysis reveals dynamic evolution of mammalian transposable elements that code for viral-like protein domains

Endogenous retroviruses (ERVs) are remnants of ancient retroviral infections of mammalian germline cells. A large proportion of ERVs lose their open reading frames (ORFs), while others retain them and become exapted by the host species. However, it remains unclear what proportion of ERVs possess ORFs (ERV-ORFs), become transcribed, and serve as candidates for co-opted genes. Hence, we investigated characteristics of 176,401 ERV-ORFs containing retroviral-like protein domains (gag, pro, pol, and env) in 19 mammalian genomes. The fractions of ERVs possessing ORFs were overall small ([~]0.15%) although they varied depending on domain types as well as species. The observed divergence of ERV-ORF from their consensus sequences suggested that a large proportion of ERV-ORFs either recently or anciently inserted themselves into mammalian genomes. Alternatively, very few ERVs lacking ORFs were found to exhibit similar divergence patterns. To identify ERV-ORFs transcribed as proteins, we compared ERV-ORFs with various multi-omics data including transcriptome data, trimethylation at histone H3 lysine 36, and transcription initiation sites from 2,834 cell types, and found 408 and 752 ERV-ORFs, accounting for 2-3% of all ERV-ORFs, with high transcriptional potential in humans and mice, respectively. Moreover, many of these ERV-ORFs with transcriptional potential were lineage-specific sequences exhibiting tissue-specific expression. These results suggest a possibility for the expression of uncharacterized functional genes containing ERV-ORFs hidden within mammalian genomes. Together, our analyses suggest that more ERV-ORFs may be co-opted in a host-species specific manner than we currently know, which are likely to have contributed to mammalian evolution and diversification.

genomics

Genome-wide 5-hydroxymethylcytosine (5hmC) emerges at early stage of in vitro hepatocyte differentiation

How cells reach different fates despite using the same DNA template, is a basic question linked to differential patterns of gene expression. Since 5-hydroxymethylcytosine (5hmC) emerged as an intermediate metabolite in active DNA demethylation, there have been increasing efforts to elucidate its function as a stable modification of the genome, including a role in establishing such tissue-specific patterns of expression. Recently we described TET1-mediated enrichment of 5hmC on the promoter region of the master regulator of hepatocyte identity, HNF4A, which precedes differentiation of liver adult progenitor cells in vitro. Here we asked whether 5hmC is involved in hepatocyte differentiation. We found a genome-wide increase of 5hmC as well as a reduction of 5-methylcytosine at early hepatocyte differentiation, a time when the liver transcript program is already established. Furthermore, we suggest that modifying s-adenosylmethionine (SAM) levels through an adenosine derivative could decrease 5hmC enrichment, triggering an impaired acquisition of hepatic identity markers. These results suggest that 5hmC is a regulator of differentiation as well as an imprint related with cell identity. Furthermore, 5hmC modulation could be a useful biomarker in conditions associated with cell de-differentiation such as liver malignancies.\n\nGraphical AbstractIt has been suggested that 5-hydroxymethylcytosine (5hmC) is an imprint of cell identity. Here we show that commitment to a hepatocyte transcriptional program is characterized by a demethylation process and emergence of 5hmC at multiple genomic locations. Cells exposed to an adenosine derivative during differentiation did not reach such 5hmC levels, and this was associated with a lower expression of hepatocyte-markers. These results suggest that 5hmC enrichment is an important step on the road to hepatocyte cell fate.\n\nO_FIG_DISPLAY_L [Figure 1] M_FIG_DISPLAY C_FIG_DISPLAY

genomics

Genome-wide association study of susceptibility to idiopathic pulmonary fibrosis

RationaleIdiopathic pulmonary fibrosis (IPF) is a complex lung disease characterised by scarring of the lung that is believed to result from an atypical response to injury of the epithelium. The mechanisms by which this arises are poorly understood and it is likely that multiple pathways are involved. The strongest genetic association with IPF is a variant in the promoter of MUC5B where each copy of the risk allele confers a five-fold risk of disease. However, genome-wide association studies have reported additional signals of association implicating multiple pathways including host defence, telomere maintenance, signalling and cell-cell adhesion.\n\nObjectivesTo improve our understanding of mechanisms that increase IPF susceptibility by identifying previously unreported genetic associations.\n\nMethods and measurementsWe performed the largest genome-wide association study undertaken for IPF susceptibility with a discovery stage comprising up to 2,668 IPF cases and 8,591 controls with replication in an additional 1,467 IPF cases and 11,874 controls. Polygenic risk scores were used to assess the collective effect of variants not reported as associated with IPF.\n\nMain resultsWe identified and replicated three new genome-wide significant (P<5x10-8) signals of association with IPF susceptibility (near KIF15, MAD1L1 and DEPTOR) and confirm associations at 11 previously reported loci. Polygenic risk score analyses showed that the combined effect of many thousands of as-yet unreported IPF risk variants contribute to IPF susceptibility.\n\nConclusionsNovel association signals support the importance of mTOR signalling in lung fibrosis and suggest a possible role of mitotic spindle-assembly genes in IPF susceptibility.

genomics

Selection following gene duplication shapes recent genome evolution in the pea aphid Acyrthosiphon pisum

Ecology of insects is as wide as their diversity, which reflects their high capacity of adaptation in most of the environments of our planet. Aphids, with over 4,000 species, have developed a series of adaptations including a high phenotypic plasticity and the ability to feed on the phloem-sap of plants, which is enriched in sugars derived from photosynthesis. Recent analyses of aphid genomes have indicated a high level of shared ancestral gene duplications that might represent a basis for genetic innovation and broad adaptations. In addition, there is a large number of recent, species-specific gene duplications whose role in adaptation remains poorly understood. Here, we tested whether duplicates specific to the pea aphid Acyrthosiphon pisum are related to genomic innovation by combining comparative genomics, transcriptomics, and chromatin accessibility analyses. Consistent with large levels of neofunctionalization, we found that most of the recent pairs of gene duplicates evolved asymmetrically, showing divergent patterns of positive selection and gene expression. Genes under selection involved a plethora of biological functions, suggesting that neofunctionalization and tissue specificity, among other evolutionary mechanisms, have orchestrated the evolution of recent paralogs in the pea aphid and may have facilitated host-symbiont cooperation. Our comprehensive phylogenomics analysis allowed to tackle the history of duplicated genes to pave the road towards understanding the role of gene duplication in ecological adaptation.

genomics

Disease Resistance Genetics and Genomics in Octoploid Strawberry

Octoploid strawberry (Fragaria x ananassa) is a valuable specialty crop, but profitable production and availability are threatened by many pathogens. Efforts to identify and introgress useful disease resistance genes (R-genes) in breeding programs are complicated by strawberrys complex octoploid genome. Recently-developed resources in strawberry, including a complete octoploid reference genome and high-resolution octoploid genotyping, enable new analyses in strawberry disease resistance genetics. This study characterizes the complete R-gene collection in the genomes of commercial octoploid strawberry and two diploid ancestral relatives, and introduces several new technological and data resources for strawberry disease resistance research. These include octoploid R-gene transcription profiling, dN/dS analysis, eQTL analysis and RenSeq analysis in cultivars. Octoploid fruit transcript expression quantitative trait loci (eQTL) were identified for 77 putative R-genes. R-genes from the ancestral diploids Fragaria vesca and Fragaria iinumae were compared, revealing differential inheritance and retention of various octoploid R-gene subtypes. The mode and magnitude of natural selection of individual F. x ananassa R-genes was also determined via dN/dS analysis. R-gene sequencing using enriched libraries (RenSeq) has been used recently for R-gene discovery in many crops, however this technique somewhat relies upon a priori knowledge of desired sequences. An octoploid strawberry capture-probe panel, derived from the results of this study, is validated in a RenSeq experiment and is presented for community use. These results give unprecedented insight into crop disease resistance genetics, and represent an advance towards exploiting variation for strawberry cultivar improvement.

genomics

Genomic architecture of human neuroanatomical diversity

Human brain anatomy is strikingly diverse and highly inheritable: genetic factors may explain up to 80% of its variability. Prior studies have tried to detect genetic variants with a large effect on neuroanatomical diversity, but those currently identified account for <5% of the variance. Here we show, based on our analyses of neuroimaging and whole-genome genotyping data from 1,765 subjects, that up to 54% of this heritability is captured by large numbers of single nucleotide polymorphisms of small effect spread throughout the genome, especially within genes and close regulatory regions. The genetic bases of neuroanatomical diversity appear to be relatively independent of those of body size (height), but shared with those of verbal intelligence scores. The study of this genomic architecture should help us better understand brain evolution and disease.

Neuroscience

Reducing INDEL calling errors in whole-genome and exome sequencing data

BackgroundINDELs, especially those disrupting protein-coding regions of the genome, have been strongly associated with human diseases. However, there are still many errors with INDEL variant calling, driven by library preparation, sequencing biases, and algorithm artifacts.\n\nMethodsWe characterized whole genome sequencing (WGS), whole exome sequencing (WES), and PCR-free sequencing data from the same samples to investigate the sources of INDEL errors. We also developed a classification scheme based on the coverage and composition to rank high and low quality INDEL calls. We performed a large-scale validation experiment on 600 loci, and find high-quality INDELs to have a substantially lower error rate than low quality INDELs (7% vs. 51%).\n\nResultsSimulation and experimental data show that assembly based callers are significantly more sensitive and robust for detecting large INDELs (>5 bp) than alignment based callers, consistent with published data. The concordance of INDEL detection between WGS and WES is low (52%), and WGS data uniquely identifies 10.8-fold more high-quality INDELs. The validation rate for WGS-specific INDELs is also much higher than that for WES-specific INDELs (85% vs. 54%), and WES misses many large INDELs. In addition, the concordance for INDEL detection between standard WGS and PCR-free sequencing is 71%, and standard WGS data uniquely identifies 6.3-fold more low-quality INDELs. Furthermore, accurate detection with Scalpel of heterozygous INDELs requires 1.2-fold higher coverage than that for homozygous INDELs. Lastly, homopolymer A/T INDELs are a major source of low-quality INDEL calls, and they are highly enriched in the WES data.\n\nConclusionsOverall, we show that accuracy of INDEL detection with WGS is much greater than WES even in the targeted region. We calculated that 60X WGS depth of coverage from the HiSeq platform is needed to recover 95% of INDELs detected by Scalpel. While this is higher than current sequencing practice, the deeper coverage may save total project costs because of the greater accuracy and sensitivity. Finally, we investigate sources of INDEL errors (e.g. capture deficiency, PCR amplification, homopolymers) with various data that will serve as a guideline to effectively reduce INDEL errors in genome sequencing.

Bioinformatics

Whole-genome sequencing is more powerful than whole-exome sequencing for detecting exome variants

We compared whole-exome sequencing (WES) and whole-genome sequencing (WGS) in six unrelated individuals. In the regions targeted by WES capture (81.5% of the consensus coding genome), the mean numbers of single-nucleotide variants (SNVs) and small insertions/deletions (indels) detected per sample were 84,192 and 13,325, respectively, for WES, and 84,968 and 12,702, respectively, for WGS. For both SNVs and indels, the distributions of coverage depth, genotype quality, and minor read ratio were more uniform for WGS than for WES. After filtering, a mean of 74,398 (95.3%) high-quality (HQ) SNVs and 9,033 (70.6%) HQ indels were called by both platforms. A mean of 105 coding HQ SNVs and 32 indels were identified exclusively by WES, whereas 692 HQ SNVs and 105 indels were identified exclusively by WGS. We Sanger sequenced a random selection of these exclusive variants. For SNVs, the proportion of false-positive variants was higher for WES (78%) than for WGS (17%). The estimated mean number of real coding SNVs (656, [~]3% of all coding HQ SNVs) identified by WGS and missed by WES was greater than the number of SNVs identified by WES and missed by WGS (26). For indels, the proportions of false-positive variants were similar for WES (44%) and WGS (46%). Finally, WES was not reliable for the detection of copy number variations, almost all of which extended beyond the targeted regions. Although currently more expensive, WGS is more powerful than WES for detecting potential disease-causing mutations within WES regions, particularly those due to SNVs.\n\nSignificanceWhole-exome sequencing (WES) is gradually being optimized to identify mutations in increasing proportions of the protein-coding exome, but whole-genome sequencing (WGS) is becoming an attractive alternative. WGS is currently more expensive than WES, but its cost should decrease more rapidly than that of WES. We compared WES and WGS on six unrelated individuals. The distribution of quality parameters for single-nucleotide variants (SNVs) and insertions/deletions (indels) was more uniform for WGS than for WES. The vast majority of SNVs and indels were identified by both techniques, but an estimated 650 high-quality coding SNVs ([~]3% of coding variants) were detected by WGS and missed by WES. WGS is therefore slightly more efficient than WES for detecting mutations in the targeted exome.

Genetics

Genome-wide comparative analysis reveals human- mouse regulatory landscape and evolution

BackgroundBecause species-specific gene expression is driven by species-specific regulation, understanding the relationship between sequence and function of the regulatory regions in different species will help elucidate how differences among species arise. Despite active experimental and computational research, the relationships among sequence, conservation, and function are still poorly understood.\n\nResultsWe compared transcription factor occupied segments (TFos) for 116 human and 35 mouse TFs in 546 human and 125 mouse cell types and tissues from the Human and the Mouse ENCODE projects. We based the map between human and mouse TFos on a one-to-one nucleotide cross-species mapper, bnMapper, that utilizes whole genome alignments (WGA).\n\nOur analysis shows that TFos are under evolutionary constraint, but a substantial portion (25.1% of mouse and 25.85% of human on average) of the TFos does not have a homologous sequence on the other species; this portion varies among cell types and TFs. Furthermore, 47.67% and 57.01% of the homologous TFos sequence shows binding activity on the other species for human and mouse respectively. However, 79.87% and 69.22% is repurposed such that it binds the same TF in different cells or different TFs in the same cells. Remarkably, within the set of TFos not showing conservation of occupancy, the corresponding genome regions in the other species are preferred locations of novel TFos. These events suggest that a substantial amount of functional regulatory sequences is exapted from other biochemically active genomic material.\n\nDespite substantial repurposing of TFos, we did not find substantial changes in their predicted target genes, suggesting that CRMs buffer evolutionary events allowing little or no change in the TF - target gene associations. Thus, the small portion of TFos with strictly conserved occupancy underestimates the degree of conservation of regulatory interactions.\n\nConclusionWe mapped regulatory sequences from an extensive number of TFs and cell types between human and mouse. A comparative analysis of this correspondence unveiled the extent of the shared regulatory sequence across TFs and cell types under study. Importantly, a large part of the shared regulatory sequence repurposed on the other species. This sequence, fueled by turnover events, provides a strong case for exaptation in regulatory elements.

Bioinformatics

Tools and Methods from the Anopheles 16 Genome Project

The dramatic reduction in sequencing costs has resulted in many initiatives to sequence certain organisms and populations. These initiatives aim to not only sequence and assemble genomes but also to perform a more broader analysis of the population structure. As part of the Anopheline Genome Consortium, which has a vested interest in studying anpopheline mosquitoes, we developed novel methods and tools to further the communities goals. We provide a brief description of these methods and tools as well as assess the contributions that each offers to the broader study of comparative genomics.

Bioinformatics