Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

Basset: Learning the regulatory code of the accessible genome with deep convolutional neural networks.

The complex language of eukaryotic gene expression remains incompletely understood. Despite the importance suggested by many noncoding variants statistically associated with human disease, nearly all such variants have unknown mechanism. Here, we address this challenge using an approach based on a recent machine learning advance--deep convolutional neural networks (CNNs). We introduce an open source package Basset (https://github.com/davek44/Basset) to apply CNNs to learn the functional activity of DNA sequences from genomics data. We trained Basset on a compendium of accessible genomic sites mapped in 164 cell types by DNaseI-seq and demonstrate far greater predictive accuracy than previous methods. Basset predictions for the change in accessibility between variant alleles were far greater for GWAS SNPs that are likely to be causal relative to nearby SNPs in linkage disequilibrium with them. With Basset, a researcher can perform a single sequencing assay in their cell type of interest and simultaneously learn that cells chromatin accessibility code and annotate every mutation in the genome with its influence on present accessibility and latent potential for accessibility. Thus, Basset offers a powerful computational approach to annotate and interpret the noncoding genome.

Genomics

The single-species metagenome: subtyping Staphylococcus aureus core genome sequences from shotgun metagenomic data

Metagenome shotgun sequence projects offer the potential for large scale biogeographic analysis of microbial species. In this project we developed a method for detecting 33 common subtypes of the pathogenic bacterium Staphylococcus aureus. We used a binomial mixture model implemented in the binstrain software and the coverage counts at > 100,000 known S. aureus SNP (single nucleotide polymorphism) sites derived from prior comparative genomic analysis to estimate the proportion of each subtype in metagenome samples. Using this pipeline we were able to obtain > 87% sensitivity and > 94% specificity when testing on low genome coverage samples of diverse S. aureus strains (0.025X). We found that 321 and 149 metagenome samples from the Human Microbiome Project and metaSUB analysis of the New York City subway, respectively, contained S. aureus at genome coverage > 0.025. In both projects, CC8 and CC30 were the most common S. aureus subtypes encountered. We found evidence that the subtype composition at different body sites of the same individual were more similar than random sampling and more limited evidence that certain body sites were enriched for particular subtypes. One surprising finding was the apparent high frequency of CC398, a lineage associated with livestock, in samples from the tongue dorsum. Epidemiologic analysis of the HMP subject population suggested that high BMI (body mass index) and health insurance are risk factors for S. aureus but there was limited power to find factors linked to carriage of even the most common subtype. In the NYC subway data, we found a small signal of geographic distance affecting subtype clustering but other unknown factors influence taxonomic distribution of the species around the city. We argue that pathogen detection in metagenome samples requires the use of subtypes based on whole species population genomic analysis rather than using ad hoc collections of reference strains.

Genomics

A high-quality reference panel reveals the complexity and distribution of structural genome changes in a human population

Structural variation (SV) represents a major source of differences between individual human genomes and has been linked to disease phenotypes. However, the majority of studies provide neither a global view of the full spectrum of these variants nor integrate them into reference panels of genetic variation.\n\nHere, we analyse whole genome sequencing data of 769 individuals from 250 Dutch families, and provide a haplotype-resolved map of 1.9 million genome variants across 9 different variant classes, including novel forms of complex indels, and retrotransposition-mediated insertions of mobile elements and processed RNAs. A large proportion are previously under reported variants sized between 21 and 100bp. We detect 4 megabases of novel sequence, encoding 11 new transcripts. Finally, we show 191 known, trait-associated SNPs to be in strong linkage disequilibrium with SVs and demonstrate that our panel facilitates accurate imputation of SVs in unrelated individuals. Our findings are essential for genome-wide association studies.

Genetics

Deep genome sequencing and variation analysis of 13 inbred mouse strains defines candidate phenotypic alleles, private variation, and homozygous truncating mutations

BackgroundThe Mouse Genomes Project is an ongoing collaborative effort to sequence the genomes of the common laboratory mouse strains. In 2011, the initial analysis of sequence variation across 17 strains found 56.7M unique SNPs and 8.8M indels. We carry out deep sequencing of 13 additional inbred strains (BUB/BnJ, C57BL/10J, C57BR/cdJ, C58/J, DBA/1J, I/LnJ, KK/HiJ, MOLF/EiJ, NZB/B1NJ, NZW/LacJ, RF/J, SEA/GnJ and ST/bJ), cataloging molecular variation within and across the strains. These strains include important models for immune response, leukemia, age-related hearing loss and rheumatoid arthritis. We now have several examples of fully sequenced closely related strains that are divergent for several disease phenotypes.\n\nResultsApproximately, 27.4M unique SNPs and 5M indels are identified across these strains compared to the C57BL/6J reference genome (GRCm38). The amount of variation found in the inbred laboratory mouse genome has increased to 71M SNPs and 12M indels. We investigate the genetic basis of highly penetrant cancer susceptibility in RF/J finding private novel missense mutations in DNA damage repair and highly cancer associated genes. We use two highly related strains (DBA/1J and DBA/2J) to investigate the genetic basis of collagen induced arthritis susceptibility.\n\nConclusionThis paper significantly expands the catalog of fully sequenced laboratory mouse strains and now contains several examples of highly genetically similar strains with divergent phenotypes. We show how studying private missense mutations can lead to insights into the genetic mechanism for a highly penetrant phenotype.

Genomics

The Next Generation Precision Medical Record - A Framework for Integrating Genomes and Wearable Sensors with Medical Records

Current medical records are rigid with regards to emerging big biomedical data. Examples of poorly integrated big data that already exist in clinical practice include whole genome sequencing and wearable sensors for real time monitoring. Genome sequencing enables conventional diagnostic interrogation and forms the fundamental baseline for precision health throughout a patients lifetime. Mobile sensors enable tailored monitoring regimes for both reducing risk through precision health interventions and acute condition surveillance. In order to address the absence of these data in the Electronic Medical Record (EMR), we worked with the SAP Personalized Medicine team to re-envision a modern medical record with these components. The pilot project used 37 patient families with complex medical records, whole genome sequencing and some level of wearable monitoring. Core functionality included patient timelines with integrated text analytics, personalized genomic curation and wearable alerts. The current phase is being rolled out to over 1500 patients in clinics across the hospital system. While fundamentally research, we believe this proof of principle platform is the first of its kind and represents the future of data driven clinical medicine.

Genomics

Copy number variants in the sheep genome detected using multiple approaches

Background.Background. Copy number variants (CNVs) are a type of polymorphism found to underlie phenotypic variation, both in humans and livestock. Most surveys of CNV in livestock have been conducted in the cattle genome, and often utilise only a single approach for the detection of copy number differences. Here we performed a study of CNV in sheep, using multiple methods to identify and characterise copy number changes. Comprehensive information from small pedigrees (trios) was collected using multiple platforms (array CGH, SNP chip and whole genome sequence data), with these data then analysed via multiple approaches to identify and verify CNVs.\n\nResults.In total, 3,488 autosomal CNV regions (CNVRs) were identified from 30 sheep. The average length of the identified CNVRs was 19kb (range of 1kb to 3.6Mb), with shorter CNVRs being more frequent than longer CNVRs. The total length of all CNVRs was 67.6Mbps, which equates to 2.7% of the sheep autosomes. For individuals this value ranged from 0.24 to 0.55%, and the majority of CNVRs were identified in single animals. Rather than being uniformly distributed throughout the genome, CNVRs tended to be clustered. Application of three independent approaches for CNVR detection facilitated a comparison of validation rates. CNVs identified on the Roche-NimbleGen 2.1M CGH array generally had low validation rates, while whole genome sequence data had the highest validation rate.\n\nConclusions.This study represents the first comprehensive survey of the distribution, prevalence and characteristics of CNVR in sheep. Multiple approaches were used to detect CNV regions and it appears that the best method for verifying CNVR on a large scale involves using a combination of detection methodologies. The characteristics of the 3,488 autosomal CNV regions identified in this study are comparable to other CNV regions reported in the literature and provide a valuable addition to the small subset of published sheep CNVs.

Genomics

The roles of LINEs, LTRs and SINEs in lineage-specific gene family expansions in the human and mouse genomes

We explored genome-wide patterns of RT content surrounding lineage-specific gene family expansions in the human and mouse genomes. Our results suggest that the size of a gene family is an important predictor of the RT distribution in close proximity to the family members. The distribution differs considerably between the three most common RT classes (LINEs, LTRs and SINEs). LINEs and LTRs tend to be more abundant around genes of multi-copy gene families, whereas SINEs tend to be depleted around such genes. Detailed analysis of the distribution and diversity of LINEs and LTRs with respect to gene family size suggests that each has a distinct involvement in gene family expansion. LTRs are associated with open chromatin sites surrounding the gene families, supporting their involvement in gene regulation, whereas LINEs may play a structural role, promoting gene duplication. This suggests that gene family expansions, especially in the mouse genome, might undergo two phases, the first is characterized by elevated deposition of LTRs and their utilization in reshaping gene regulatory networks. The second phase is characterized by rapid gene family expansion due to continuous accumulation of LINEs and it appears that, in some instances at least, this could become a runaway process. We provide an example in which this has happened and we present a simulation supporting the possibility of the runaway process. Our observations also suggest that specific differences exist in this gene family expansion process between human and mouse genomes.

Genomics

Whole Genome Analysis of 132 Clinical Saccharomyces cerevisiae Strains Reveals Extensive Ploidy Variation

Budding yeast has undergone several independent transitions from commercial to clinical lifestyles. The frequency of such transitions suggests that clinical yeast strains are derived from environmentally available yeast populations, including commercial sources. However, despite their important role in adaptive evolution, the prevalence of polyploidy and aneuploidy has not extensively analyzed in clinical strains. In this study, we have looked for patterns governing the transition to clinical invasion in the largest screen of clinical yeast isolates to date. In particular, we have focused on the hypothesis that ploidy changes have influenced adaptive processes. We sequenced 145 yeast strains, 132 of which are clinical isolates. We found pervasive large-scale genomic variation in both overall ploidy (34% of strains identified as 3n/4n) and individual chromosomal copy numbers (36% of strains identified as aneuploid). We also found evidence for the highly dynamic nature of yeasts genomes, with 35 strains showing partial chromosomal copy number changes and 8 strains showing multiple independent chromosomal events. Intriguingly, a lineage identified to be baker/commercial derived with a unique damaging mutation in NDC80 was particularly prone to polyploidy, with 83% of its members being triploid or tetraploid. Polyploidy was in turn associated with a >2x increase in aneuploidy rates as compared to other lineages. This dataset provides a rich source of information of the genomics of clinical yeast strains and highlights the potential importance of large-scale genomic copy variation in yeast adaptation.

Genomics

Inferring heterozygosity from ancient and low coverage genomes

While genetic diversity can be quantified accurately from high coverage sequencing, it is often desirable to obtain such estimates from low coverage data, either to save costs or because of low DNA quality as observed for ancient samples. Here we introduce a method to accurately infer heterozygosity probabilistically from very low coverage sequences of a single individual. The method relaxes the infinite sites assumption of previous methods, does not require a reference sequence and takes into account both variable sequencing errors and potential post-mortem damage. It is thus also applicable to non-model organisms and ancient genomes. Since error rates as reported by sequencing machines are generally distorted and require recalibration, we also introduce a method to infer accurately recalibration parameter in the presence of post-mortem damage. This method does also not require knowledge about the underlying genome sequence, but instead works from haploid data (e.g. from the X-chromosome from mammalian males) and integrates over the unknown genotypes. Using extensive simulations we show that a few Mb of haploid data is sufficient for accurate recalibration even at average coverages as low as 1-3x. At similar coverages, out method also produces very accurate estimates of heterozygosity down to 10-4 within windows of about 1Mb. We further illustrate the usefulness of our approach by inferring genome-wide patterns of diversity for several ancient human samples and found that 3,000-5,000 samples showed diversity patterns comparable to modern humans. In contrast, two European hunter-gatherer samples exhibited not only considerably lower levels of diversity than modern samples, but also highly distinct distributions of diversity along their genomes. Interestingly, these distributions were also very differently between the two samples, supporting earlier conclusions of a highly diverse and structured population in Europe prior to the arrival of farming.

Genomics

Eukaryotic Association Module in Phage WO Genomes from Wolbachia

Viruses are trifurcated into eukaryotic, archaeal and bacterial categories. This domain-specific ecology underscores why eukaryotic genes are typically co-opted by eukaryotic viruses and bacterial genes are commonly found in bacteriophages. However, the presence of bacteriophages in symbiotic bacteria that obligately reside in eukaryotes may promote eukayotic DNA transfers to bacteriophages. By sequencing full genomes from purified bacteriophage WO particles of Wolbachia, we discover a novel eukaryotic association module with various animal proteins domains, such as the black widow latrotoxin-CTD, that are uninterrupted in intact bacteriophage genomes, enriched with eukaryotic protease cleavage sites, and combined with additional domains to forge some of the largest bacteriophage genes (up to 14,256 bp). These various protein domain families are central to eukaryotic functions and have never before been reported in packaged bacteriophages, and their phylogeny, distribution and sequence diversity implies lateral transfer from animal to bacteriophage genomes. We suggest that the evolution of these eukaryotic protein domains in bacteriophage WO parallels the evolution of eukaryotic genes in canonical eukaryotic viruses, namely those commandeered for viral life cycle adaptations. Analogous selective pressures and evolutionary outcomes may occur in bacteriophage WO as a result of its \"two-fold cell challenge\" to persist in and traverse cells of obligate intracellular bacteria that strictly reside in animal cells. Finally, the full WO genome sequences and identification of attachment sites will advance eventual genetic manipulation of Wolbachia for disease control strategies.

Genomics

Prospects of genomic prediction in the USDA Soybean Germplasm Collection: Historical data creates robust models for enhancing selection of accessions.

The identification and mobilization of useful genetic variation from germplasm banks for use in breeding programs is critical for future genetic gain and protection against crop pests. Plummeting costs of next-generation sequencing and genotyping is revolutionizing the way in which researchers and breeders interface with plant germplasm collections. An example of this is the high density genotyping of the entire USDA Soybean Germplasm Collection. We assessed the usefulness of 50K SNP data collected on 18,480 domesticated soybean (G. max) accessions and vast historical phenotypic data for developing genomic prediction models for protein, oil, and yield. Resulting genomic prediction models explained an appreciable amount of the variation in accession performance in independent validation trials, with correlations between predicted and observed reaching up to 0.92 for oil and protein and 0.79 for yield. The optimization of training set design was explored using a series of cross-validation schemes. It was found that the target population and environment need to be well represented in the training set. Secondly, genomic prediction training sets appear to be robust to the presence of data from diverse geographical locations and genetic clusters. This finding, however, depends on the influence of shattering and lodging, and may be specific to soybean with its presence of maturity groups. The distribution of 7,608 non-phenotyped accessions was examined through the application of genomic prediction models. The distribution of predictions of phenotyped accessions was representative of the distribution of predictions for non-phenotyped accessions, with no non-phenotyped accessions being predicted to fall far outside the range of predictions of phenotyped accessions.

Genomics

Genome resources for climate-resilient cowpea, an essential crop for food security

Cowpea (Vigna unguiculata L. Walp.) is a legume crop that is resilient to hot and drought-prone climates, and a primary source of protein in sub-Saharan Africa and other parts of the developing world. However, genome resources for cowpea have lagged behind most other major crop plants. Here we describe foundational genome resources and their application to analysis of germplasm currently in use in West African breeding programs. Resources developed from the African cultivar IT97K-499-35 include bacterial artificial chromosome (BAC) libraries and a BAC-based physical map, assembled sequences from 4,355 BACs, as well as a whole-genome shotgun (WGS) assembly. These resources and WGS sequences of an additional 36 diverse cowpea accessions supported the development of a genotyping assay for over 50,000 SNPs, which was then applied to five biparental RIL populations to produce a consensus genetic map containing 37,372 SNPs. This genetic map enabled the anchoring of 100 Mb of WGS and 420 Mb of BAC sequences, an exploration of genetic diversity along each linkage group, and clarification of macrosynteny between cowpea and common bean. The genomes of West African breeding lines and landraces have regions of marked depletion of diversity, some of which coincide with QTL that may be the result of artificial selection or environmental adaptation. The new publicly available resources and knowledge help to define goals and accelerate the breeding of improved varieties to address food security issues related to limited-input small-holder farming and climate stress.

Genomics

Elucidating the genetic basis of an oligogenic birth defect using whole genome sequence data in a non-model organism, Bubalus bubalis

Recent strong selection for dairy traits in water buffalo has been associated with higher levels of inbreeding, leading to an increase in the prevalence of genetic diseases such as transverse hemimelia (TH), a congenital developmental abnormality characterized by the absence of a variable distal portion of the hindlimbs. The limited genomic resources available for water buffalo, in conjunction with an unconfirmed inheritance pattern, required an original approach to identify genetic variants associated with this disease. The genomes of 4 bilaterally affected cases, 7 unilaterally affected cases, and 14 controls were sequenced. Variant calling identified 19.8 million high confidence single nucleotide polymorphisms (SNPs) and 2.8 million insertions/deletions (INDELs). A concordance analysis of SNPs and INDELs requiring all unilateral and bilateral cases and none of the controls to be homozygous for the same allele, revealed two genes, WNT7A and SMARCA4, known to play a role in embryonic hindlimb development. Additionally, SNP alleles in NOTCH1 and RARB were homozygous exclusively in the bilaterally affected cases, suggesting an oligogenic mode of inheritance. Homozygosity mapping by whole genome de novo assembly was then used to identify large contigs representing regions of homozygosity in the cases. This also supported an oligogenic mode of inheritance; implicating 13 genes involved in aberrant hindlimb development in the bilateral cases and 11 in the unilateral cases. A genome-wide association study (GWAS) predicted additional modifier genes. Results from these analyses suggest that mutations in SMARCA4 and WNT7A are required for expression of TH, while several other loci including NOTCH1 act as modifiers and increase the severity of the disease phenotype. Although our data show that the inheritance of TH is complex, we predict that homozygous variants in WNT7A and SMARCA4 are necessary for the expression of TH and selection against these variants and avoidance of carrier-to-carrier matings should eradicate TH.\n\nAuthor SummaryGenetic diseases often occur and are spread through small populations under strong selection where rates of inbreeding can be significant. The use of a limited number of water buffalo males via artificial insemination for genetic improvement of milk and milk composition has increased the frequency of the genetic disease, transverse hemimelia (TH). Transverse hemimelia affected calves are normally developed except for malformation of one or both hindlimbs or both hindlimbs and one or both forelimbs. Little is known about the inheritance pattern of TH. We discovered genetic variants present in cases where both hindlimbs and one forelimb were affected, cases were both hindlimbs were affected, cases where only one hindlimb was affected, and in non-affected water buffalo that predict TH to be inherited as an oligogenic disease with two driver loci necessary for disease expression and several additional modifier genes that are responsible for the severity of the disease phenotype. We predict that selection against mutations in the two major loci and the avoidance of mating animals that are heterozygous for these mutations will eliminate TH from water buffalo.

Genomics

MINOTAUR: A platform for the analysis and visualization of multivariate results from genome scans with R Shiny

Genome scans are widely used to identify 'outliers' in genomic data: loci with different patterns compared with the rest of the genome due to the action of selection or other non-adaptive forces of evolution. These genomic datasets are often high-dimensional, with complex correlation structures among variables, making it a challenge to identify outliers in a robust way. The Mahalanobis distance has been widely used for this purpose, but has the major limitation of assuming that data follow a simple parametric distribution. Here we develop three new metrics that can be used to identify outliers in multivariate space, while making no strong assumptions about the distribution of the data. These metrics are implemented in the R package MINOTAUR, which also includes an interactive web-based application for visualizing outliers in high-dimensional datasets. We illustrate how these metrics can be used to identify outliers from simulated genetic data, and discuss some of the limitations they may face in application.

Genomics

The queenslandensis and the type form of the dengue fever mosquito (Aedes aegypti L.) are genomically indistinguishable

BackgroundThe mosquito Aedes aegypti (L.) is a major vector of viral diseases like dengue fever, Zika and chikungunya. Aedes aegypti exhibits high morphological and behavioral variation, some of which is thought to be of epidemiological significance. Globally distributed domestic Ae. aegypti have been traditionally grouped into (i) the very pale variety queenslandensis and (ii) the type form. Because the two color forms co-occur across most of their range, there is interest in understanding how freely they interbreed. This knowledge is particularly important for control strategies that rely on mating compatibilities between the release and target mosquitoes, such as Wolbachia releases and SIT. To answer this question, we analyzed nuclear and mitochondrial genome-wide variation in the co-occurring pale and type Ae. aegypti from northern Queensland (Australia) and Singapore.\n\nMethods/FindingsWe typed 74 individuals at a 1170 bp-long mitochondrial sequence and at 16,569 nuclear SNPs using a customized double-digest RAD sequencing. 11/29 genotyped individuals from Singapore and 11/45 from Queensland were identified as var. queenslandensis based on the diagnostic scaling patterns. We found 24 different mitochondrial haplotypes, seven of which were shared between the two forms. Multivariate genetic clustering based on nuclear SNPs corresponded to individuals geographic location, not their color. Several family groups consisted of both forms and three queenslandensis individuals were Wolbachia infected, indicating previous breeding with the type form which has been used to introduce Wolbachia into Ae. aegypti populations.\n\nConclusionAedes aegypti queenslandensis are genomically indistinguishable from the type form, which points to these forms freely interbreeding at least in Australia and Singapore. Based on our findings, it is unlikely that the presence of very pale Ae. aegypti will affect the success of Aedes control programs based on Wolbachia-infected, sterile or RIDL mosquitoes.\n\nAuthor SummaryAedes aegypti, the most important vector of dengue and Zika, greatly varies in body color and behavior. Two domestic forms of this mosquito, the very pale queenslandensis and the browner type, are often found together in populations around the globe. Knowing how freely they interbreed is important for the control strategies such as releases of Wolbachia and sterile males. To answer this question, we used RAD sequencing to genotype samples of both forms collected in Singapore and northern Queensland. We did not find any association between the mitochondrial or nuclear genome-wide variation and color variation in these populations. Rather, \"paleness\" is likely to be a quantitative trait under some environmental influence. We also detected several queenslandensis individuals with the Wolbachia infection, indicating free interbreeding with the type form which has been used to introduce Wolbachia into Ae. aegypti populations. Overall, our data show that the very pale queenslandensis are not genomically separate, and their presence is unlikely to affect the success of Aedes control programs based on Wolbachia-infected, sterile or RIDL mosquitoes.

Genomics

SMRT Genome Assembly Corrects Reference Errors, Resolving the Genetic Basis of Virulence in Mycobacterium tuberculosis

The genetic basis of virulence in Mycobacterium tuberculosis has been investigated through genome comparisons of its virulent (H37Rv) and attenuated (H37Ra) sister strains. Such analysis, however, relies heavily on the accuracy of the sequences. While the H37Rv reference genome has had several corrections to date, that of H37Ra is unmodified since its original publication. Here, we report the assembly and finishing of the H37Ra genome from single-molecule, real-time (SMRT) sequencing. Our assembly reveals that the number of H37Ra-specific variants is less than half of what the Sanger-based H37Ra reference sequence indicates, undermining and, in some cases, invalidating the conclusions of several studies. PE_PPE family genes, which are intractable to commonly-used sequencing platforms because of their repetitive and GC-rich nature, are overrepresented in the set of genes in which all reported H37Ra-specific variants are contradicted. We discuss how our results change the picture of virulence attenuation and the power of SMRT sequencing for producing high-quality reference genomes.

Genomics

Rewired RNAi-Mediated Genome Surveillance in House Dust Mites

House dust mites are common pests with an unusual evolutionary history, being descendants of a parasitic ancestor. Transition to parasitism is frequently accompanied by genome rearrangements, possibly to accommodate the genetic change needed to access new ecology. Transposable element (TE) activity is a source of genomic instability that can trigger large-scale genomic alterations. Eukaryotes have multiple transposon control mechanisms, one of which is RNA interference (RNAi). Investigation of the dust mite genome failed to identify a major RNAi pathway: the Piwi-associated RNA (piRNA) pathway, which has been replaced by a novel small-interfering RNAs (siRNAs)-like pathway. Co-opting of piRNA function by dust mite siRNAs is extensive, including establishment of TE control master loci that produce siRNAs. Interestingly, other members of the Acari have piRNAs indicating loss of this mechanism in dust mites is a recent event. Flux of RNAi-mediated control of TEs provides a mechanism for unusual arc of dust mite evolution.

Genomics

GSuite HyperBrowser: integrative analysis of dataset collections across the genome and epigenome

Genome-wide, cell-type-specific profiles are being systematically generated for numerous genomic and epigenomic features. There is, however, no universally applicable analytical methodology for such data. We present GSuite HyperBrowser, the first comprehensive solution for integrative analysis of dataset collections across the genome and epigenome. The GSuite HyperBrowser is an open-source system for streamlined acquisition and customizable statistical analysis of large collections of genome-wide datasets. The system is based on new computational and statistical methodologies that permit comparative and confirmatory analyses across multiple disparate data sources. Expert guidance and reproducibility are facilitated via a Galaxy-based web-interface. The software is available at https://hyperbrowser.uio.no/gsuite

Genomics