Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Rapid Core-Genome Alignment and Visualization for Thousands of Intraspecific Microbial Genomes

Though many microbial species or clades now have hundreds of sequenced genomes, existing whole-genome alignment methods do not efficiently handle comparisons on this scale. Here we present the Harvest suite of core-genome alignment and visualization tools for quickly analyzing thousands of intraspecific microbial strains. Harvest includes Parsnp, a fast core-genome multi-aligner, and Gingr, a dynamic visual platform. Combined they provide interactive core-genome alignments, variant calls, recombination detection, and phylogenetic trees. Using simulated and real data we demonstrate that our approach exhibits unrivaled speed while maintaining the accuracy of existing methods. The Harvest suite is open-source and freely available from: http://github.com/marbl/harvest.

Bioinformatics

annotatr: Associating genomic regions with genomic annotations

MotivationAnalysis of next-generation sequencing data often results in a list of genomic regions. These may include differentially methylated CpGs/regions, transcription factor binding sites, interacting chromatin regions, or GWAS-associated SNPs, among others. A common analysis step is to annotate such genomic regions to genomic annotations (promoters, exons, enhancers, etc.). Existing tools are limited by a lack of annotation sources and flexible options, the time it takes to annotate regions, an artificial one-to-one region-to-annotation mapping, a lack of visualization options to easily summarize data, or some combination thereof.\n\nResultsWe developed the annotatr Bioconductor package to flexibly and quickly summarize and plot annotations of genomic regions. The annotatr package reports all intersections of regions and annotations, giving a better understanding of the genomic context of the regions. A variety of graphics functions are implemented to easily plot numerical or categorical data associated with the regions across the annotations, and across annotation intersections, providing insight into how characteristics of the regions differ across the annotations. We demonstrate that annotatr is up to 27x faster than comparable R packages. Overall, annotatr enables a richer biological interpretation of experiments.\n\nAvailabilityhttp://bioconductor.org/packages/annotatr/\n\nContactrcavalca@umich.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

Bioinformatics

Genome Graphs and the Evolution of Genome Inference

The human reference genome is part of the foundation of modern human biology, and a monumental scientific achievement. However, because it excludes a great deal of common human variation, it introduces a pervasive reference bias into the field of human genomics. To reduce this bias, it makes sense to draw on representative collections of human genomes, brought together into reference cohorts. There are a number of techniques to represent and organize data gleaned from these cohorts, many using ideas implicitly or explicitly borrowed from graph based models. Here, we survey various projects underway to build and apply these graph based structures--which we collectively refer to as genome graphs--and discuss the improvements in read mapping, variant calling, and haplotype determination that genome graphs are expected to produce.

bioinformatics

The 3D Genome Browser: a web-based browser for visualizing 3D genome organization and long-range chromatin interactions

Recent advent of 3C-based technologies such as Hi-C and ChIA-PET provides us an opportunity to explore chromatin interactions and 3D genome organization in an unprecedented scale and resolution. However, it remains a challenge to visualize chromatin interaction data due to its size and complexity. Here, we introduce the 3D Genome Browser (http://3dgenome.org), which allows users to conveniently explore both publicly available and their own chromatin interaction data. Users can also seamlessly integrate other \"omics\" data sets, such as ChIP-Seq and RNA-Seq for the same genomic region, to gain a complete view of both regulatory landscape and 3D genome structure for any given gene. Finally, our browser provides multiple methods to link distal cis-regulatory elements with their potential target genes, including virtual 4C, ChIA-PET, Capture Hi-C and cross-cell-type correlation of proximal and distal DNA hypersensitive sites, and therefore represents a valuable resource for the study of gene regulation in mammalian genomes.

bioinformatics

Evolution of gene expression after whole-genome duplication: new insights from the spotted gar genome

Whole genome duplications (WGD) are important evolutionary events. Our understanding of underlying mechanisms, including the evolution of duplicated genes after WGD, however remains incomplete. Teleost fish experienced a common WGD (teleost-specific genome duplication, or TGD) followed by a dramatic adaptive radiation leading to more than half of all vertebrate species. The analysis of gene expression patterns following TGD at the genome level has been limited by the lack of suitable genomic resources. The recent concomitant release of the genome sequence of spotted gar (a representative of holosteans, the closest lineage of teleosts that lacks the TGD) and the tissue-specific gene expression repertoires of over 20 holostean and teleostean fish species, including spotted gar, zebrafish and medaka (the PhyloFish project), offered a unique opportunity to study the evolution of gene expression following TGD in teleosts. We show that most TGD duplicates gained their current status (loss of one duplicate gene or retention of both duplicates) relatively rapidly after TGD (i.e. prior to the divergence of medaka and zebrafish lineages). The loss of one duplicate is the most common fate after TGD with a probability of approximately 80%. In addition, the fate of duplicate genes after TGD, including subfunctionalization, neofunctionalization, or retention of two similar copies occurred not only before, but also after the radiation of species tested, in consistency with a role of the TGD in speciation and/or evolution of gene function. Finally, we report novel cases of TGD ohnolog subfunctionalization and neofunctionalization that further illustrate the importance of these processes.

evolutionary biology

Genome-wide association mapping and genomic prediction unravels CBSD resistance in a Manihot esculenta breeding population

Cassava (Manihot esculenta Crantz), a key carbohydrate dietary source for millions of people in Africa, faces severe yield loses due to two viral diseases: cassava brown streak disease (CBSD) and cassava mosaic disease (CMD). The completion of the cassava genome sequence and the whole genome marker profiling of clones from African breeding programs (www.nextgencassava.org) provides cassava breeders the opportunity to deploy additional breeding strategies and develop superior varieties with both farmer and industry preferred traits. Here the identification of genomic segments associated with resistance to CBSD foliar symptoms and root necrosis as measured in two breeding panels at different growth stages and locations is reported. Using genome-wide association mapping and genomic prediction models we describe the genetic architecture for CBSD severity and identify loci strongly associated on chromosomes 4 and 11. Moreover, the significantly associated region on chromosome 4 colocalises with a Manihot glaziovii introgression segment and the significant SNP markers on chromosome 11 are situated within a cluster of nucleotide-binding site leucine-rich repeat (NBS-LRR) genes previously described in cassava. Overall, predictive accuracy values found in this study varied between CBSD severity traits and across GS models with Random Forest and RKHS showing the highest predictive accuracies for foliar and root CBSD severity scores.

genetics

GenoGAM 2.0: Scalable and efficient implementation of genome-wide generalized additive models for gigabase-scale genomes

GenoGAM (Genome-wide generalized additive models) is a powerful statistical modeling tool for the analysis of ChIP-Seq data with flexible factorial design experiments. However large runtime and memory requirements of its current implementation prohibit its application to gigabase-scale genomes such as mammalian genomes.\n\nHere we present GenoGAM 2.0 a scalable and efficient implementation that is 2-3 orders of magnitude faster than the previous version. This is achieved by exploiting the sparsity of the model using the SuperLU direct solver for parameter fitting, and sparse Cholesky factorization together with the sparse inverse subset algorithm for computing standard errors. Furthermore the HDF5 library is employed to store data efficiently on hard drive, reducing memory footprint while keeping I/O low. Whole-genome fits for human ChIP-seq datasets (ca. 300 million parameters) could be obtained in less than 9 hours on a standard 60-core server. GenoGAM 2.0 is implemented as an open source R package and currently available on GitHub. A Bioconductor release of the new version is in preparation.\n\nWe have vastly improved the performance of the GenoGAM framework, opening up its application to all types of organisms. Moreover, our algorithmic improvements for fitting large GAMs could be of interest to the statistical community beyond the genomics field.

bioinformatics

Genome of octoploid plant maca (Lepidium meyenii) illuminates genomic basis for high altitude adaptation in the central Andes

Maca (Lepidium meyenii Walp, 2n = 8x = 64) of Brassicaceae family is an Andean economic plant cultivated on the 4000-4500 meters central sierra in Peru. Considering the rapid uplift of central Andes occurred 5 to 10 million years ago (Mya), an evolutionary question arises on how plants like maca acquire high altitude adaptation within short geological period. Here, we report the high-quality genome assembly of maca, in which two close-spaced maca-specific whole genome duplications (WGDs, [~] 6.7 Mya) were identified. Comparative genomics between maca and close-related Brassicaceae species revealed expansions of maca genes and gene families involved in abiotic stress response, hormone signaling pathway and secondary metabolite biosynthesis via WGDs. Retention and subsequent evolution of many duplicated genes may account for the morphological and physiological changes (i.e. small leaf shape and loss of vernalization) in maca for high altitude environment. Additionally, some duplicated maca genes under positive selection were identified with functions in morphological adaptation (i.e. MYB59) and development (i.e. GDPD5 and HDA9). Collectively, the octoploid maca genome sheds light on the important roles of WGDs in plant high altitude adaptation in the Andes.

Genomics

Improved genome assembly of American alligator genome reveals conserved architecture of estrogen signaling

The American alligator, Alligator mississippiensis, like all crocodilians, has temperature-dependent sex determination, in which the sex of an embryo is determined by the incubation temperature of the egg during a critical period of development. The lack of genetic differences between male and female alligators leaves open the question of how the genes responsible for sex determination and differentiation are regulated. One insight into this question comes from the fact that exposing an embryo incubated at male-producing temperature to estrogen causes it to develop ovaries. Because estrogen response elements are known to regulate genes over long distances, a contiguous genome assembly is crucial for predicting and understanding its impact.\n\nWe present an improved assembly of the American alligator genome, scaffolded with in vitro proximity ligation (Chicago) data. We use this assembly to scaffold two other crocodilian genomes based on synteny. We perform RNA sequencing of tissues from American alligator embryos to find genes that are differentially expressed between embryos incubated at male-versus female-producing temperature. Finally, we use the improved contiguity of our assembly along with the current model of CTCF-mediated chromatin looping to predict regions of the genome likely to contain estrogen-responsive genes. We find that these regions are significantly enriched for genes with female-biased expression in developing gonads after the critical period during which sex is determined by incubation temperature. We thus conclude that estrogen signaling is a major driver of female-biased gene expression in the post-temperature sensitive period gonads.

Genomics

Genome-wide association study of HIV whole genome sequences validated using drug resistance

Genome-wide association studies (GWAS) have considerably advanced our understanding of human traits and diseases. With the increasing availability of whole genome sequences (WGS) for pathogens, it is important to establish whether GWAS of viral genomes could reveal important biological insights. Here we perform the first proof of concept viral GWAS examining drug resistance (DR), a phenotype with well understood genetics.\n\nWe performed a GWAS of DR in a sample of 343 HIV subtype C patients failing 1st line antiretroviral treatment in rural KwaZulu-Natal, South Africa. The majority and minority variants within each sequence were called using PILON, and GWAS was performed within PLINK. HIV WGS from patients failing on different antiretroviral treatments were compared to sequences derived from individuals naive to the respective treatment.\n\nGWAS methodology was validated by identifying five associations on a genetic level that led to amino acid changes known to cause DR. Further, we highlighted the ability of GWAS to identify epistatic effects, identifying two replicable variants within amino acid 68 of the reverse transcriptase protein previously described as potential fitness compensatory mutations. A possible additional DR variant within amino acid 91 of the matrix region of the Gag protein was associated with tenofovir failure, highlighting the ability of GWAS to identify variants outside classical candidate genes. Our results also suggest a polygenic component to DR.\n\nThese results validate the applicability of GWAS to HIV WGS data even in relative small samples, and emphasise how high throughput sequencing can provide novel and clinically relevant insights. Further they suggested that for viruses like HIV, population structure was only minor concern compared to that seen in bacteria or parasite GWAS. Given the small genome length and reduced burden for multiple testing, this makes HIV an ideal candidate for GWAS.

Genomics

The house spider genome reveals an ancient whole-genome duplication during arachnid evolution

The duplication of genes can occur through various mechanisms and is thought to make a major contribution to the evolutionary diversification of organisms. There is increasing evidence for a large-scale duplication of genes in some chelicerate lineages including two rounds of whole genome duplication (WGD) in horseshoe crabs. To investigate this further we sequenced and analyzed the genome of the common house spider Parasteatoda tepidariorum. We found pervasive duplication of both coding and non-coding genes in this spider, including two clusters of Hox genes. Analysis of synteny conservation across the P. tepidariorum genome suggests that there has been an ancient WGD in spiders. Comparison with the genomes of other chelicerates, including that of the newly sequenced bark scorpion Centruroides sculpturatus, suggests that this event occurred in the common ancestor of spiders and scorpions and is probably independent of the WGDs in horseshoe crabs. Furthermore, characterization of the sequence and expression of the Hox paralogs in P. tepidariorum suggests that many have been subject to neofunctionalization and/or subfunctionalization since their duplication, and therefore may have contributed to the diversification of spiders and other pulmonate arachnids.

genomics

Myotis rufoniger Genome Sequence and Analyses: M. rufoniger’s Genomic Feature and the Decreasing Effective Population Size of Myotis Bats

Myotis rufoniger is a vesper bat in the genus Myotis. Here we report the whole genome sequence and analyses of the M. rufoniger. We generated 124 Gb of short-read DNA sequences with an estimated genome size of 1.88 Gb at a sequencing depth of 66x fold. The sequences were aligned to M. brandtii bat reference genome at a mapping rate of 96.50% covering 95.71% coding sequence region at 10x coverage. The divergence time of Myotis bat family is estimated to be 11.5 million years, and the divergence time between M. rufoniger and its closest species M. davidii is estimated to be 10.4 million years. We found 1,239 function-altering M. rufoniger specific amino acid sequences from 929 genes compared to other Myotis bat and mammalian genomes. The functional enrichment test of the 929 genes detected amino acid changes in melanin associated DCT, SLC45A2, TYRP1, and OCA2 genes possibly responsible for the M. rufonigers red fur color and a general coloration in Myotis. N6AMT1 gene, associated with arsenic resistance, showed a high degree of function alteration in M. rufoniger. We further confirmed that M. rufoniger also has bat-specific sequences within FSHB, GHR, IGF1R, TP53, MDM2, SLC45A2, RGS7BP, RHO, OPN1SW, and CNGB3 genes that have already been published to be related to bats reproduction, lifespan, flight, low vision, and echolocation. Additionally, our demographic history analysis found that the effective population size of Myotis clade has been consistently decreasing since [~]30k years ago. M. rufonigers effective population size was the lowest in Myotis bats, confirming its relatively low genetic diversity.

genomics

The grayling genome reveals selection on gene expression regulation after whole genome duplication

Whole genome duplication (WGD) has been a major evolutionary driver of increased genomic complexity in vertebrates. One such event occurred in the salmonid family ~80 million years ago (Ss4R) giving rise to a plethora of structural and regulatory duplicate-driven divergence, making salmonids an exemplary system to investigate the evolutionary consequences of WGD. Here, we present a draft genome assembly of European grayling (Thymallus thymallus) and use this in a comparative framework to study evolution of gene regulation following WGD. Among the Ss4R duplicates identified in European grayling and Atlantic salmon (Salmo salar), one third reflect non-neutral tissue expression evolution, with strong purifying selection, maintained over ~50 million years. Of these, the majority reflect conserved tissue regulation under strong selective constraints related to brain and neural-related functions, as well as higher-order protein-protein interactions. A small subset of the duplicates has evolved tissue regulatory expression divergence in a common ancestor, which have been subsequently conserved in both lineages, suggestive of adaptive divergence following WGD. These candidates for adaptive tissue expression divergence have elevated rates of protein coding- and promoter-sequence evolution and are enriched for immune- and lipid metabolism ontology terms. Lastly, lineage-specific duplicate divergence points towards underlying differences in adaptive pressures on expression regulation in the non-anadromous grayling versus the anadromous Atlantic salmon.\n\nOur findings enhance our understanding of the role of WGD in genome evolution and highlights cases of regulatory divergence of Ss4R duplicates, possibly related to a niche shift in early salmonid evolution.

genomics

Whole-genome sequencing analysis of genomic copy number variation (CNV) using low-coverage and paired-end strategies is highly efficient and outperforms array based CNV analysis

BackgroundCNV analysis is an integral component to the study of human genomes in both research and clinical settings. Array-based CNV analysis is the current first-tier approach in clinical cytogenetics. Decreasing costs in high-throughput sequencing and cloud computing have opened doors for the development of sequencing-based CNV analysis pipelines with fast turnaround times. We carry out a systematic and quantitative comparative analysis for several low-coverage whole-genome sequencing (WGS) strategies to detect CNV in the human genome.\n\nMethodsWe compared the CNV detection capabilities of WGS strategies (short-insert, 3kb-, and 5kb-insert mate-pair) each at 1x, 3x, and 5x coverages relative to each other and to 17 currently used high-density oligonucleotide arrays. For benchmarking, we used a set of Gold Standard (GS) CNVs generated for the 1000-Genomes-Project CEU subject NA12878.\n\nResultsOverall, low-coverage WGS strategies detect drastically more GS CNVs compared to arrays and are accompanied with smaller percentages of CNV calls without validation. Furthermore, we show that WGS (at [≥]1x coverage) is able to detect all seven GS deletion-CNVs >100 kb in NA12878 whereas only one is detected by most arrays. Lastly, we show that the much larger 15 Mbp Cri-du-chat deletion can be readily detected with short-insert paired-end WGS at even just 1x coverage.\n\nConclusionsCNV analysis using low-coverage WGS is efficient and outperforms the array-based analysis that is currently used for clinical cytogenetics.

genomics

A model species for agricultural pest genomics: the genome of the Colorado potato beetle, Leptinotarsa decemlineata (Coleoptera: Chrysomelidae)

The Colorado potato beetle is one of the most challenging agricultural pests to manage. It has shown a spectacular ability to adapt to a variety of solanaceaeous plants and variable climates during its global invasion, and, notably, to rapidly evolve insecticide resistance. To examine evidence of rapid evolutionary change, and to understand the genetic basis of herbivory and insecticide resistance, we tested for structural and functional genomic changes relative to other arthropod species using genome sequencing, transcriptomics, and community annotation. Two factors that might facilitate rapid evolutionary change include transposable elements, which comprise at least 17% of the genome and are rapidly evolving compared to other Coleoptera, and high levels of nucleotide diversity in rapidly growing pest populations. Adaptations to plant feeding are evident in gene expansions and differential expression of digestive enzymes in gut tissues, as well as expansions of gustatory receptors for bitter tasting. Surprisingly, the suite of genes involved in insecticide resistance is similar to other beetles. Finally, duplications in the RNAi pathway might explain why Leptinotarsa decemlineata has high sensitivity to dsRNA. The L. decemlineata genome provides opportunities to investigate a broad range of phenotypes and to develop sustainable methods to control this widely successful pest.

genomics

The Medical Genome Reference Bank: a whole-genome data resource of 4,000 healthy elderly individuals. Rationale and cohort design

Allele frequency data from human reference populations is of increasing value for filtering and assignment of pathogenicity to genetic variants. Aged and healthy populations are more likely to be selectively depleted of pathogenic alleles, and therefore particularly suitable as a reference populations for the major diseases of clinical and public health importance. However, reference studies of the healthy elderly have remained under-represented in human genetics. We have developed the Medical Genome Reference Bank (MGRB), a large-scale comprehensive whole-genome dataset of confirmed healthy elderly individuals, to provide a publicly accessible resource for health-related research, and for clinical genetics. It also represents a useful resource for studying the genetics of healthy aging. The MGRB comprises 4,000 healthy, older individuals with no reported history of cancer, cardiovascular disease or dementia, recruited from two Australian community-based cohorts. DNA derived from blood samples will be subject to whole genome sequencing. The MGRB will measure genome-wide genetic variation in 4,000 individuals, mostly of European decent, aged 60-95 years (mean age [≥] 75 years). The MGRB has committed to a policy of data sharing, employing a hierarchical data management system to maintain participant privacy and confidentiality, whilst maximizing research and clinical usage of the database. The MGRB will represent a dataset of international significance, broadly accessible to the clinical and genetic research community.

genomics

IDENTIFICATION OF GENOMIC REGIONS CARRYING A CAUSAL MUTATION IN UNORDERED GENOMES

Whole genome sequencing using high-throughput sequencing (HTS) technologies offers powerful opportunities to study genetic variation. Mapping the mutations responsible for different phenotypes is generally an involved and time-consuming process so researchers have developed user-friendly tools for mapping-by-sequencing, yet they are not applicable to organisms with non-sequenced genomes. We introduce SDM (SNP Distribution Method), a reference independent method for rapid discovery of mutagen-induced mutations in typical forward genetic screens. SDM aims to order a disordered collection of HTS reads or contigs such that the fragment carrying the causative mutation can be identified. SDM uses typical distributions of homozygous SNPs that are linked to a phenotype-altering SNP in a non-recombinant region as a model to order the fragments. To implement and test SDM, we created model genomes with an idealised SNP density based on Arabidopsis thaliana chromosome 1 and analysed fragments with size distribution similar to reads or contigs assembled from HTS sequencing experiments. SDM groups the contigs by their normalised SNP density and arranges them to maximise the fit to the expected SNP distribution. We tested the procedure in existing datasets by examining SNP distributions in recent out-cross and back-cross experiments in Arabidopsis thaliana backgrounds. In all the examples we analysed, homozygous SNPs were normally distributed around the causal mutation. We used the real SNP densities obtained from these experiments to prove the efficiency and accuracy of SDM. The algorithm was able to successfully identify small sized (10-100 kb) genomic regions containing the causative mutation.

Bioinformatics

Genome-wide expression profiling Drosophila melanogaster deficiency heterozygotes reveals diverse genomic responses.

Deletions, commonly referred to as deficiencies by Drosophila geneticists, are valuable tools for mapping genes and for genetic pathway discovery via dose-dependent suppressor and enhancer screens. More recently, it has become clear that deviations from normal gene dosage are associated with multiple disorders in a range of species including humans. While we are beginning to understand some of the transcriptional effects brought about by gene dosage changes and the chromosome rearrangement breakpoints associated with them, much of this work relies on isolated examples. We have systematically examined deficiencies on the left arm of chromosome 2 and characterize gene-by-gene dosage responses that vary from collapsed expression through modest partial dosage compensation to full or even over compensation. We found negligible long-range effects of creating novel chromosome domains at deletion breakpoints, suggesting that cases of changes in gene regulation due to altered nuclear architecture are rare. These rare cases include trans de-repression when deficiencies delete chromatin characterized as repressive in other studies. Generally, effects of breakpoints on expression are promoter proximal (~100 bp) or within the gene body. Genome-wide effects of deficiencies are observed at genes with regulatory relationships to genes within the deleted segments, highlighting the subtle expression network defects in these sensitized genetic backgrounds.\n\nAuthor summaryDeletions alter gene dose in heterozygotes and bring distant regions of the genome into juxtaposition. We find that the transcriptional dose response is generally varied, gene-specific, and coherently propagates into gene expression regulatory networks. Analysis of deletion heterozygote expression profiles indicates that distinct genetic pathways are weakened in adult flies bearing different deletions even though they show minimal or no overt phenotypes. While there are exceptions, breakpoints have a minimal effect on the expression of flanking genes, despite the fact that different regions of the genome are brought into contact and that important elements such as insulators are deleted. These data suggest that there is little effect of nuclear architecture and long-range enhancer and/or silencer promoter contact on gene expression in the compact Drosophila genome.

Genetics