Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,603 records · Page 89Linked to original sources

CIDER-Seq: unbiased virus enrichment and single-read, full length genome sequencing

Deep-sequencing of virus isolates using short-read sequencing technologies is problematic since viruses are often present in complexes sharing a high-degree of sequence identity. The full-length genomes of such highly-similar viruses cannot be assembled accurately from short sequencing reads. We present a new method, CIDER-Seq (Circular DNA Enrichment Sequencing) which successfully generates accurate full-length virus genomes from individual sequencing reads with no sequence assembly required. CIDER-Seq operates by combining a PCR-free, circular DNA enrichment protocol with Single Molecule Real Time sequencing and a new sequence deconcatenation algorithm. We apply our technique to produce more than 1,200 full-length, highly accurate geminivirus genomes from RNAi-transgenic and control plants in a field trial in Kenya. Using CIDER-Seq we can demonstrate for the first time that the expression of antiviral doublestranded RNA (dsRNA) in transgenic plants causes a consistent shift in virus populations towards species sharing low homology to the transgene derived dsRNA. Our results show that CIDER-seq is a powerful, cost-effective tool for accurately sequencing circular DNA viruses, with future applications in deep-sequencing other forms of circular DNA such as transposons and plasmids.

genomics

N6-Methyladenine DNA Modification in Human Genome

DNA N6-methyladenine (6mA) modification is the most prevalent DNA modification in prokaryotes, but whether it exists in human cells and whether it plays a role in human diseases remain enigmatic. Here, we showed that 6mA is extensively present in human genome, and we cataloged 881,240 6mA sites accounting for [~]0.051% of the total adenines. [G/C]AGG[C/T] was the most significantly associated motif with 6mA modification. 6mA sites were enriched in the coding regions and mark actively transcribed genes in human cells. We further found that DNA N6-methyladenine and N6-demethyladenine modification in human genome were mediated by methyltransferase N6AMT1 and demethylase ALKBH1, respectively. The abundance of 6mA was significantly lower in cancers, accompaning with decreased N6AMT1 and increased ALKBH1 levels, and down-regulation of 6mA modification levels promoted tumorigenesis. Collectively, our results demonstrate that DNA 6mA modification is extensively present in human cells and the decrease of genomic DNA 6mA promotes human tumorigenesis.

genomics

Spatial Genome Organization as a Framework for Somatic Alterations in Human Cancer

Genomic material within the nucleus is folded into successive layers in order to package and organize the long string of linear DNA. This hierarchical level of folding is closely associated with transcriptional regulation and DNA replication. Microscopic studies concluded that genome organization inside the nucleus is not random and chromosomes within the nucleus create territories1. More recent chromosome conformation studies have revealed that mammalian chromosomes are structured into tissue-invariant topologically associating domains (TADs) where the DNA within a given domain interacts more frequently together than with regions in other domains2,3. Genes within the same TADs represent similar expression and histone-modification profiles4. In addition, enhancer-promoter pairs within the same TAD respond similarly to hormone induction5 or differentiation cues6. Therefore, regions separating different TADs (boundaries) have important roles in reinforcing the stability of these domain-wide features. Indeed, TAD boundary disruptions in human genetic disorders7,8 or human cancers lead to misregulation of certain genes9,10, due to de novo enhancer exposure to promoters. Here, to understand effects and distributions of somatic structural variations across TADs, we utilized single nucleotide variations, deletions, inversions, tandem-duplications and complex rearrangements from 2658 high-coverage whole genome sequencing data across various cancer types with paired normal samples. We comprehensively profiled structural variations with respect to their effect on TAD boundaries, on the regulation of genes in human cancers.

genomics

Piggy: A Rapid, Large-Scale Pan-Genome Analysis Tool for Intergenic Regions in Bacteria

Despite overwhelming evidence that variation in intergenic regions (IGRs) in bacteria impacts on phenotypes, most current approaches for analysing pan-genomes focus exclusively on protein-coding sequences. To address this we present Piggy, a novel pipeline that emulates Roary except that it is based only on IGRs. We demonstrate the use of Piggy for pan-genome analyses of Staphylococcus aureus and Escherichia coli using large genome datasets. For S. aureus, we show that highly divergent (\"switched\") IGRs are associated with differences in gene expression, and we establish a multi-locus reference database of IGR alleles (igMLST; implemented in BIGSdb). Piggy is available at https://github.com/harry-thorpe/piggy.

genomics

No major flaws in "Identification of individuals by trait prediction using whole-genome sequencing data"

In a recently published PNAS article, we studied the identifiability of genomic samples using machine learning methods [Lippert et al., 2017]. In a response, Erlich [2017] argued that our work contained major flaws. The main technical critique of Erlich [2017] builds on a simulation experiment that shows that our proposed algorithm, which uses only a genomic sample for identification, performed no better than a strategy that uses demographic variables. Below, we show why this comparison is misleading and provide a detailed discussion of the key critical points in our analyses that have been brought up in Erlich [2017] and in the media. Further, not only faces may be derived from DNA, but a wide range of phenotypes and demographic variables. In this light, the main contribution of Lippert et al. [2017] is an algorithm that identifies genomes of individuals by combining multiple DNA-based predictive models for a myriad of traits.

genomics

A survey of DNA methylation polymorphism identifies environmentally responsive co-regulated networks of epigenetic variation in the human genome

While studies such as the 1000 Genomes Projects have resulted in detailed maps of genetic variation in humans, to date there are few robust maps of epigenetic variation. We defined sites of common epigenetic variation, termed Variably Methylated Regions (VMRs) in five purified cell types. We observed that VMRs occur preferentially at enhancers and 3 UTRs. While the majority of VMRs have high heritability, a subset of VMRs within the genome show highly correlated variation in trans, forming co-regulated networks that have low heritability, differ between cell types and are enriched for specific transcription factor binding sites and biological pathways of functional relevance to each tissue. For example, in T cells we defined a network of 72 co-regulated VMRs enriched for genes with roles in T-cell activation; in fibroblasts a network of 21 coregulated VMRs comprising all four HOX gene clusters enriched for control of tissue growth; and in neurons a network of 112 VMRs enriched for roles in learning and memory. By culturing genetically-identical fibroblasts under varying conditions of nutrient deprivation and cell density, we experimentally demonstrate that some VMR networks are responsive to environmental conditions, with methylation levels at these loci changing in a coordinated fashion in trans dependent on cellular growth. Intriguingly these environmentally-responsive VMRs showed a strong enrichment for imprinted loci (p<10-94), suggesting that these are particularly sensitive to environmental conditions. Our study provides a detailed map of common epigenetic variation in the human genome, showing that both genetic and environmental causes underlie this variation.

genomics

Integrated analysis sheds light on evolutionary trajectories of young transcription start sites in the human genome

Previous studies revealed widespread transcription initiation and fast turnover of transcription start sites (TSSs) in mammalian genomes. Yet how new TSSs originate and how they evolve over time remain poorly understood. To address these questions, we analyzed [~]200,000 human TSSs by integrating evolutionary and functional genomic data, particularly focusing on TSSs that emerged in the primate lineages. We found that intrinsic factors of repetitive sequences and their proximity to established regulatory modules (extrinsic factors) contribute significantly to origin of new TSSs. In early periods, young TSSs experience rapid sequence evolution driven by endogenous mutational mechanisms that reduce the instability of associated repetitive sequences. In later periods, the regulatory functions of young TSSs are gradually modified, and with evolutionary changes subject to temporal (fewer regulatory changes in younger TSSs) and spatial constraints (fewer regulatory changes in more isolated TSSs). These findings advance our understanding of how regulatory innovations arise in the genome throughout evolution and highlight the roles of repetitive sequences in these processes.

genomics

High-resolution genome-wide functional dissection of transcriptional regulatory regions in human

Genome-wide epigenomic maps revealed millions of regions showing signatures of enhancers, promoters, and other gene-regulatory elements1. However, high-throughput experimental validation of their function and high-resolution dissection of their driver nucleotides remain limited in their scale and length of regions tested. Here, we present a new method, HiDRA (High-Definition Reporter Assay), that overcomes these limitations by combining components of Sharpr-MPRA2 and STARR-Seq3 with genome-wide selection of accessible regions from ATAC-Seq4. We used HiDRA to test ~7 million DNA fragments preferentially selected from accessible chromatin in the GM12878 lymphoblastoid cell line. By design, accessibility-selected fragments were highly overlapping (up to 370 per region), enabling us to pinpoint driver regulatory nucleotides by exploiting subtle differences in reporter activity between partially-overlapping fragments, using a new machine learning model SHARPR2. Our resulting maps include ~65,000 regions showing significant enhancer function and enriched for endogenous active histone marks (including H3K9ac, H3K27ac), regulatory sequence motifs, and regions bound by immune regulators. Within them, we discover ~13,000 high-resolution driver elements enriched for regulatory motifs and evolutionarily-conservednucleotides, and help predict causal genetic variants underlying disease from genome-wide association studies. Overall, HiDRA provides a general, scalable, high-throughput, and high-resolution approach for experimental dissection of regulatory regions and driver nucleotides in the context of human biology and disease.

genomics

Improved Prokaryotic Gene Prediction Yields Insights into Transcription and Translation Mechanisms on Whole Genome Scale

In a conventional view of the prokaryotic genome organization promoters precede operons and RBS sites with Shine-Dalgarno consensus precede genes. However, recent experimental research suggesting a more diverse view motivated us to develop an algorithm with improved gene-finding accuracy. We describe GeneMarkS-2, an ab initio algorithm that uses a model derived by self-training for finding species-specific (native) genes, along with an array of pre-computed heuristic models designed to identify harder-to-detect genes (likely horizontally transferred). Importantly, we designed GeneMarkS-2 to identify several types of distinct sequence patterns (signals) involved in gene expression control, among them the patterns characteristic for leaderless transcription as well as non-canonical RBS patterns. To assess the accuracy of GeneMarkS-2 we used genes validated by COG annotation, proteomics experiments, and N-terminal protein sequencing. We observed that GeneMarkS-2 performed better on average in all accuracy measures when compared with the current state-of-the-art gene prediction tools. Furthermore, the screening of [~]5,000 representative prokaryotic genomes made by GeneMarkS-2 predicted frequent leaderless transcription in both archaea and bacteria. We also observed that the RBS sites in some species with leadered transcription did not necessarily exhibit the Shine-Dalgarno consensus. The modeling of different types of sequence motifs regulating gene expression prompted a division of prokaryotic genomes into five categories with distinct sequence patterns around the gene starts.\n\n[Supplemental material is available for this article].

genomics

The genome sequence of the soft-rot fungus Penicillium purpurogenum reveals a high gene dosage for lignocellulolytic enzymes

The high lignocellulolytic activity displayed by the soft-rot fungus P. purpurogenum has made it a target for the study of novel lignocellulolytic enzymes. We have obtained a reference genome of 36.2Mb of non-redundant sequence (11,057 protein-coding genes). The 49 largest scaffolds cover 90% of the assembly, and CEGMA analysis reveals that our assembly covers most if not all all protein-coding genes. RNASeq was performed and 93.1% of the reads aligned within the assembled genome. These data, plus the independent sequencing of a set of genes of lignocellulose-degrading enzymes, validate the quality of the genome sequence. P. purpurogenum shows a higher number of proteins with CAZy motifs, transcription factors and transporters as compared to other sequenced Penicillia. These results demonstrate the great potential for lignocellulolytic activity of this fungus and the possible use of its enzymes in related industrial applications.

genomics

Landscape genomic prediction for restoration of a Eucalyptus foundation species under climate change

As species face rapid environmental change, we can build resilient populations through restoration projects that incorporate predicted future climates into seed sourcing decisions. Eucalyptus melliodora is a foundation species of a critically endangered community in Australia that is a target for restoration. We examined patterns of genomic and phenotypic variation to make empirical based recommendations for seed sourcing. We examined isolation by distance and isolation by environment, determining gene flow up to 500 km and associations with environmental variables. Climate chamber studies revealed extensive phenotypic variation both within and among sampling sites, but no site-specific differentiation in phenotypic plasticity. Overall our results suggest that seed can be sourced broadly across the landscape, providing ample diversity for adaptation to environmental change. Application of our landscape genomic model to E. melliodora restoration projects can identify genomic variation suitable for predicted future climates, thereby increasing the long term probability of successful restoration.

genomics

Whole genome sequence and comparative analysis of Borrelia burgdorferi MM1

Lyme disease is caused by spirochaetes of the Borrelia burgdorferi sensu lato genospecies. Complete genome assemblies are available for fewer than ten strains of Borrelia burgdorferi sensu stricto, the primary cause of Lyme disease in North America. MM1 is a sensu stricto strain originally isolated in the midwestern United States. Aside from a small number of genes, the complete genome sequence of this strain has not been reported. Here we present the complete genome sequence of MM1 in relation to other sensu stricto strains and in terms of its Multi Locus Sequence Typing. Our results indicate that MM1 is a new sequence type which contains a conserved main chromosome and 15 plasmids. Our results include the first contiguous 28.5 kb assembly of lp28-8, a linear plasmid carrying the vls antigenic variation system, from a Borrelia burgdorferi sensu stricto strain.

genomics

Parental allele-specific genome architecture and transcription during the cell cycle

A normal human somatic cell inherits two haploid genomes. Individual chromosomes of each pair have distinct parental origins and parental alleles are known to unequally contribute to cellular function. We integrated chromosome conformation (form) and gene transcription (function) analyses to dissect the dynamics of the maternal and paternal genomes in lymphoblastoid cells during the cell cycle. We found a distinct set of homologous alleles with very different activity often located close to boundaries of euchromatin and heterochromatin domains. We also identified a set of allele-biased topologically associating domains (TADs) that were small sized and had higher gene density. Thousands of genes show allelically biased expression (ABE) with false discovery rate < 0.05, and 98% of them have no allelic switching during G1, S, and G2/M phases. A subset of ABE genes are preferentially localized near TAD boundaries, enriched with chromatin organization transcription factor binding sites, and contained higher number of sequence variants in CCCTC-binding factor sites. Our results extend previous findings of sequence variation as a basis for unequal functional parental genomes. Investigation of haplotype-resolved form-function dynamics may further our understanding of phenotypic traits, genetic diseases, vulnerability to complex disorders, and the development of cancers.

genomics

Comparative genomics of Mycobacterium africanum Lineage 5 and Lineage 6 from Ghana suggests different ecological niches

Mycobacterium africanum (Maf) causes up to half of human tuberculosis in West Africa, but little is known on this pathogen. We compared the genomes of 253 Maf clinical isolates from Ghana, including both L5 and L6. We found that the genomic diversity of L6 was higher than in L5, and the selection pressures differed between both groups. Regulatory proteins appeared to evolve neutrally in L5 but under purifying selection in L6. Conversely, human T cell epitopes were under purifying selection in L5, but under positive selection in L6. Although only 10% of the T cell epitopes were variable, mutations were mostly lineage-specific. Our findings indicate that Maf L5 and L6 are genomically distinct, possibly reflecting different ecological niches.

genomics

Germline Contamination and Leakage in Whole Genome Somatic Single Nucleotide Variant Detection

BackgroundThe clinical sequencing of cancer genomes to personalize therapy is becoming routine across the world. However, concerns over patient re-identification from these data lead to questions about how tightly access should be controlled. It is not thought to be possible to re-identify patients from somatic variant data. However, somatic variant detection pipelines can mistakenly identify germline variants as somatic ones, a process called \"germline leakage\". The rate of germline leakage across different somatic variant detection pipelines is not well-understood, and it is uncertain whether or not somatic variant calls should be considered re-identifiable. To fill this gap, we quantified germline leakage across 259 sets of whole-genome somatic single nucleotide variant (SNVs) predictions made by 21 teams as part of the ICGC-TCGA DREAM Somatic Mutation Calling Challenge.\n\nResultsThe median somatic SNV prediction set contained 4,325 somatic SNVs and leaked one germline polymorphism. The level of germline leakage was inversely correlated with somatic SNV prediction accuracy and positively correlated with the amount of infiltrating normal cells. The specific germline variants leaked differed by tumour and algorithm. To aid in quantitation and correction of leakage, we created a tool, called GermlineFilter, for use in public-facing somatic SNV databases.\n\nConclusionsThe potential for patient re-identification from leaked germline variants in somatic SNV predictions has led to divergent open data access policies, based on different assessments of the risks. Indeed, a single, well-publicized re-identification event could reshape public perceptions of the values of genomic data sharing. We find that modern somatic SNV prediction pipelines have low germline-leakage rates, which can be further reduced, especially for cloud-sharing, using pre-filtering software.

genomics

BoostMe accurately predicts DNA methylation values in whole-genome bisulfite sequencing of multiple human tissues

BackgroundBisulfite sequencing is widely employed to study the role of DNA methylation in disease; however, the data suffer from biases due to variability in depth of coverage. Imputation of methylation values at low-coverage sites may mitigate these biases while also identifying important genomic features and motifs associated with predictive power.\n\nResultsHere we describe BoostMe, a novel method for imputation of DNA methylation within whole-genome bisulfite sequencing (WGBS) data based on a gradient boosting algorithm. Importantly, we designed a new feature that leverages information from multiple samples in the same tissue and disease state, enabling BoostMe to outperform existing imputation methods in speed and accuracy. We show that imputation improves WGBS concordance with the Infinium MethylationEPIC array at low WGBS sequencing depth, suggesting improvement in WGBS accuracy after imputation. Furthermore, we compare the ability of BoostMe and DeepCpG - a deep neural network method - to identify interesting features and motifs associated with methylation in three human tissues implicated in type 2 diabetes (T2D) etiology. We find that while BoostMe only identifies features important to general methylation levels across tissues, DeepCpG is able to learn differences in methylation-associated sequence motifs among different tissues and identify tissue-specific regulators of differentiation such as EBF1 in adipose, ASCL2 in muscle, and FOXA1, TCF12, and NRF1 in islets. Neither algorithm readily identified T2D-associated features.\n\nConclusionsOur findings demonstrate the current power and limitations of machine and deep learning algorithms to both improve the quality of, and infer biological meaning from, genome-wide DNA methylation data.

genomics

Germline determinants of the somatic mutation landscape in 2,642 cancer genomes

Cancers develop through somatic mutagenesis, however germline genetic variation can markedly contribute to tumorigenesis via diverse mechanisms. We discovered and phased 88 million germline single nucleotide variants, short insertions/deletions, and large structural variants in whole genomes from 2,642 cancer patients, and employed this genomic resource to study genetic determinants of somatic mutagenesis across 39 cancer types. Our analyses implicate damaging germline variants in a variety of cancer predisposition and DNA damage response genes with specific somatic mutation patterns. Mutations in the MBD4 DNA glycosylase gene showed association with elevated C>T mutagenesis at CpG dinucleotides, a ubiquitous mutational process acting across tissues. Analysis of somatic structural variation exposed complex rearrangement patterns, involving cycles of templated insertions and tandem duplications, in BRCA1-deficient tumours. Genome-wide association analysis implicated common genetic variation at the APOBEC3 gene cluster with reduced basal levels of somatic mutagenesis attributable to APOBEC cytidine deaminases across cancer types. We further inferred over a hundred polymorphic L1/LINE elements with somatic retrotransposition activity in cancer. Our study highlights the major impact of rare and common germline variants on mutational landscapes in cancer.

genomics

Lamins organize the global three-dimensional genome from the nuclear periphery

Lamins are structural components of the nuclear lamina (NL) that regulate genome organization and gene expression, but the mechanism remains unclear. Using Hi-C, we show that lamins maintain proper interactions among the topologically associated chromatin domains (TADs) but not their overall architecture. Combining Hi-C with fluorescence in situ hybridization (FISH) and analyses of lamina-associated domains (LADs), we reveal that lamin loss causes expansion or detachment of specific LADs in mouse ES cells. The detached LADs disrupt 3D interactions of both LADs and interior chromatin. 4C and epigenome analyses further demonstrate that lamins maintain the active and repressive chromatin domains among different TADs. By combining these studies with transcriptome analyses, we found a significant correlation between transcription changes and the changes of active and inactive chromatin domain interactions. These findings provide a foundation to further study how the nuclear periphery impacts genome organization and transcription in development and NL-associated diseases.\n\nHighlightsO_LILamin loss does not affect the overall TAD structure but alters TAD-TAD interactions\nC_LIO_LILamin null ES cells exhibit decondensation or detachment of specific LAD regions\nC_LIO_LIExpansion and detachment of LADs can alter genome-wide 3D chromatin interactions\nC_LIO_LIAltered chromatin domain interactions are correlated with altered transcription\nC_LI

genomics