Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,063 records · Page 59Linked to original sources

Modular dynamics of DNA co-methylation networks exposes the functional organization of colon cancer cells’ genome

Epigenomic plasticity is interconnected with chromatin structure and gene regulation. In tumor progression, orchestrated remodeling of genome organization accompanies the acquisition of malignant properties. DNA methylation, a key epigenetic mark extensively altered in cancer, is also linked to genome architecture and function. Based on this association, we postulate that the dissection of long-range co-methylation structure unveils cancer cells genome architecture remodeling. We applied network-modeling of DNA methylation co-variation in two colon cancer cohorts and found abundant and consistent transchromosomal structures in both normal and tumor tissue. Normal-tumor comparison indicated substantial remodeling of the epigenome covariation and revealed novel genomic compartments with a unique signature of DNA methylation rank inversion.

genomics

Leveraging evolutionary relationships to improve Anopheles genome assemblies

While new sequencing technologies have lowered financial barriers to whole genome sequencing, resulting assemblies are often fragmented and far from finished. Subsequent improvements towards chromosomal-level status can be achieved by both experimental and computational approaches. Requiring only annotated assemblies and gene orthology data, comparative genomics approaches that aim to capture evolutionary signals to predict scaffold neighbours (adjacencies) offer potentially substantive improvements without the costs associated with experimental scaffolding or re-sequencing. We leverage the combined detection power of three such gene synteny-based methods applied to 21 Anopheles mosquito assemblies with variable contiguity levels to produce consensus sets of scaffold adjacency predictions. Three complementary validations were performed on subsets of assemblies with additional supporting data: six with physical mapping data; 13 with paired-end RNA sequencing (RNAseq) data; and three with new assemblies based on re-scaffolding or incorporating Pacific Biosciences (PacBio) sequencing data. Improved assemblies were built by integrating the consensus adjacency predictions with supporting experimental data, resulting in 20 new reference assemblies with improved contiguities. Combined with physical mapping data for six anophelines, chromosomal positioning of scaffolds improved assembly anchoring by 47% for A. funestus and 38% A. stephensi. Reconciling an A. funestus PacBio assembly with synteny-based and RNAseq-based adjacencies and physical mapping data resulted in a new 81.5% chromosomally mapped reference assembly and cytogenetic photomap. While complementary experimental data are clearly key to achieving high-quality chromosomal-level assemblies, our assessments and validations of gene synteny-based computational methods highlight the utility of applying comparative genomics approaches to improve community genomic resources.

genomics

Comparison of single-cell whole-genome amplification strategies

Single-cell genomics is an alluring area that holds the potential to change the way we understand cell populations. Due to the small amount of DNA within a single cell, whole-genome amplification becomes a mandatory step in many single-cell applications. Unfortunately, single-cell whole-genome amplification (scWGA) strategies suffer from several technical biases that complicate the posterior interpretation of the data. Here we compared the performance of six different scWGA methods (GenomiPhi, REPLIg, TruePrime, Ampli1, MALBAC, and PicoPLEX) after amplifying and low-pass sequencing the complete genome of 230 healthy/tumoral human cells. Overall, REPLIg outperformed competing methods regarding DNA yield, amplicon size, amplification breadth, amplification uniformity -being the only method with a random amplification bias-, and false single-nucleotide variant calls. On the other hand, non-MDA methods, and in particular Ampli1, showed less allelic imbalance and ADO, more reliable copy-number profiles and less chimeric amplicons. While no single scWGA method showed optimal performance for every aspect, they clearly have distinct advantages. Our results provide a convenient guide for selecting a scWGA method depending on the question of interest while revealing relevant weaknesses that should be considered during the analysis and interpretation of single-cell sequencing data.

genomics

Draft genome assembly and population genetics of an agricultural pollinator, the solitary alkali bee (Halictidae: Nomia melanderi)

Alkali bees (Nomia melanderi) are solitary relatives of the halictine bees, which have become an important model for the evolution of social behavior, but for which few solitary comparisons exist. These ground-nesting bees defend their developing offspring against pathogens and predators, and thus exhibit some of the key traits that preceded insect sociality. Alkali bees are also efficient native pollinators of alfalfa seed, which is a crop of major economic value in the United States. We sequenced, assembled, and annotated a high-quality draft genome of 299.6 Mbp for this species. Repetitive content makes up more than one-third of this genome, and previously uncharacterized transposable elements are the most abundant type of repetitive DNA. We predicted 10,847 protein coding genes, and identify 479 of these undergoing positive directional selection with the use of population genetic analysis based on low-coverage whole genome sequencing of 19 individuals. We found evidence of recent population bottlenecks, but no significant evidence of population structure. We also identify 45 genes enriched for protein translation and folding, transcriptional regulation, and triglyceride metabolism evolving slower in alkali bees compared to other halictid bees. These resources will be useful for future studies of bee comparative genomics and pollinator health research.

genomics

A Draft Genome and High-Density Genetic Mapof European Hazelnut (Corylus avellana L.)

European hazelnut (Corylus avellana L.) is of global agricultural and economic significance, with genetic diversity existing in hundreds of accessions. Breeding efforts have focused on maximizing nut yield and quality and reducing susceptibility to diseases such as Eastern filbert blight (EFB). Here we present the first sequenced genome among the order Fagales, the EFB-resistant diploid hazelnut accession Jefferson (OSU 703.007). We assembled the highly heterozygous hazelnut genome using an Illumina only approach and the final assembly has a scaffold N50 of 21.5kb. We captured approximately 91 percent (345 Mb) of the flow-cytometry-determined genome size and identified 34,910 putative gene loci. In addition, we identified over 2 million polymorphisms across seven diverse hazelnut accessions and characterized t heir effect on coding sequences. We produced t wo high-density genetic maps with 3,209 markers from an F1 hazelnut population, representing a five-fold increase in marker density over previous maps. These genomic resources will aide in the discovery of molecular markers linked to genes of interest for hazelnut breeding efforts, and are available to the community at https://www.cavellanagenomeportal.com/.

genomics

Scaling computational genomics to millions of individuals with GPUs

Current genomics methods were designed to handle tens to thousands of samples, but will soon need to scale to millions to keep up with the pace of data and hypothesis generation in biomedical science. Moreover, costs associated with processing these growing datasets will become prohibitive without improving the computational efficiency and scalability of methods. Here, we show that recently developed machine-learning libraries (TensorFlow and PyTorch) facilitate implementation of genomics methods for GPUs and significantly accelerate computations. To demonstrate this, we re-implemented methods for two commonly performed computational genomics tasks: QTL mapping and Bayesian non-negative matrix factorization. Our implementations ran > 200 times faster than current CPU-based versions, and these analyses are [~]5-10 fold cheaper on GPUs due to the vastly shorter runtimes. We anticipate that the accessibility of these libraries, and the improvements in run-time will lead to a transition to GPU-based implementations for a wide range of computational genomics methods.

genomics

CNApp: a web-based tool for integrative analysis of genomic copy number alterations in cancer

Somatic copy number alterations (CNAs) are a hallmark of cancer. Although CNA profiles have been established for most human tumor types, their precise role in tumorigenesis as well as their clinical and therapeutic relevance remain largely unclear. Thus, computational and statistical approaches are required to thoroughly define the interplay between CNAs and tumor phenotypes. Here we developed CNApp, a user-friendly web tool that offers sample- and cohort-level computational analyses, allowing a comprehensive and integrative exploration of CNAs with clinical and molecular variables. By using purity-corrected segmented data from multiple genomic platforms, CNApp generates genome-wide profiles, computes CNA scores for broad, focal and global CNA burdens, and uses machine learning-based predictions to classify samples. We applied CNApp to a pan-cancer dataset of 10,635 genomes from TCGA showing that CNA patterns classify cancer types according to their tissue-of-origin, and that broad and focal CNA scores positively correlate in samples with low amounts of whole-chromosome and chromosomal arm-level imbalances. Moreover, using the hepatocellular carcinoma cohort from the TCGA repository, we demonstrate the reliability of the tool in identifying recurrent CNAs, confirming previous results. Finally, we establish machine learning-based models to predict colon cancer molecular subtypes and microsatellite instability based on broad CNA scores and specific genomic imbalances. In summary, CNApp facilitates data-driven research and provides a unique framework for the first time to comprehensively assess CNAs and perform integrative analyses that enable the identification of relevant clinical implications. CNApp is hosted at http://cnapp.bsc.es.

genomics

Genome-wide association study of anti-Mullerian hormone levels in pre-menopausal women of late reproductive age and relationship with genetic determinants of reproductive lifespan

Anti-Mullerian hormone (AMH) is required for sexual differentiation in the fetus, and in adult females AMH is produced by growing ovarian follicles. Consequently, AMH levels are correlated with ovarian reserve, declining towards menopause when the oocyte pool is exhausted. A previous genome-wide association study identified three genetic variants in and around the AMH gene that explained 25% of variation in AMH levels in adolescent males but did not identify any genetic associations reaching genome-wide significance in adolescent females. To explore the role of genetic variation in determining AMH levels in women of late reproductive age, we carried out a genome-wide meta-analysis in 3,344 pre-menopausal women from five cohorts (median age 44-48 years at blood draw). A single genetic variant, rs16991615, previously associated with age at menopause, reached genome-wide significance at P=3.48x10-10, with a per allele difference in age-adjusted inverse normal AMH of 0.26 SD (95% CI [0.18,0.34]). We investigated whether genetic determinants of female reproductive lifespan were more generally associated with pre-menopausal AMH levels. Genetically-predicted age at menarche had no robust association but genetically-predicted age at menopause was associated with lower AMH levels by 0.18 SD (95% CI [0.14,0.21]) in age-adjusted inverse normal AMH per one-year earlier age at menopause. Our findings support the hypothesis that AMH is a valid measure of ovarian reserve in pre-menopausal women and suggest that the underlying biology of ovarian reserve results in a causal link between pre-menopausal AMH levels and menopause timing.

genomics

Generating closed bacterial genomes from long-read nanopore sequencing of microbiomes

We present the first method for efficient recovery of complete, closed genomes directly from microbiomes using nanopore long-read sequencing and assembly. We apply our approach to three healthy human gut communities and compare results to short read and read cloud approaches. We obtain nine finished genomes including the first reported closed genome of Prevotella copri, an organism with highly repetitive genome structure prevalent in non-western human gut microbiomes.

genomics

Genome-wide association study of suicide attempt in psychiatric disorders identifies association with major depression polygenic risk scores

ObjectiveOver 90% of suicide attempters have a psychiatric diagnosis, however twin and family studies suggest that the genetic etiology of suicide attempt (SA) is partially distinct from that of the psychiatric disorders themselves. Here, we present the largest genome-wide association study (GWAS) on suicide attempt using major depressive disorder (MDD), bipolar disorder (BIP) and schizophrenia (SCZ) cohorts from the Psychiatric Genomics Consortium.\n\nMethodSamples comprise 1622 suicide attempters and 8786 non-attempters with MDD, 3264 attempters and 5500 non-attempters with BIP and 1683 attempters and 2946 non-attempters with SCZ. SA GWAS were performed comparing attempters to non-attempters in each disorder followed by meta-analysis across disorders. Polygenic risk scoring investigated the genetic relationship between SA and the psychiatric disorders.\n\nResultsThree genome-wide significant loci for SA were found: one associated with SA in MDD, one in BIP, and one in the meta-analysis of SA in mood disorders. These associations were not replicated in independent mood disorder cohorts from the UK Biobank and iPSYCH. Polygenic risk scores for major depression were significantly associated with SA in MDD (P=0.0002), BIP (P=0.0006) and SCZ (P=0.0006).\n\nConclusionsThis study provides new information on genetic associations and the genetic etiology of SA across psychiatric disorders. The finding that polygenic risk scores for major depression predict suicide attempt across disorders provides a possible starting point for predictive modelling and preventative strategies. Further collaborative efforts to increase sample size hold potential to robustly identify genetic associations and gain biological insights into the etiology of suicide attempt.

genetics

Population genomics of parallel hybrid zones in the mimetic butterflies, H. melpomene and H. erato

Hybrid zones can be valuable tools for studying evolution and identifying genomic regions responsible for adaptive divergence and underlying phenotypic variation. Hybrid zones between subspecies of Heliconius butterflies can be very narrow and are maintained by strong selection acting on colour pattern. The co-mimetic species H. erato and H. melpomene have parallel hybrid zones where both species undergo a change from one colour pattern form to another. We use restriction associated DNA sequencing to obtain several thousand genome wide sequence markers and use these to analyse patterns of population divergence across two pairs of parallel hybrid zones in Peru and Ecuador. We compare two approaches for analysis of this type of data; alignment to a reference genome and de novo assembly, and find that alignment gives the best results for species both closely (H. melpomene) and distantly (H. erato, ~15% divergent) related to the reference sequence. Our results confirm that the colour pattern controlling loci account for the majority of divergent regions across the genome, but we also detect other divergent regions apparently unlinked to colour pattern differences. We also use association mapping to identify previously unmapped colour pattern loci, in particular the Ro locus. Finally, we identify within our sample a new cryptic population of H. timareta in Ecuador, which occurs at relatively low altitude and is mimetic with H. melpomene malleti.

Evolutionary Biology

Species Delimitation using Genome-Wide SNP Data

The multi-species coalescent has provided important progress for evolutionary inferences, including increasing the statistical rigor and objectivity of comparisons among competing species delimitation models. However, Bayesian species delimitation methods typically require brute force integration over gene trees via Markov chain Monte Carlo (MCMC), which introduces a large computation burden and precludes their application to genomic-scale data. Here we combine a recently introduced dynamic programming algorithm for estimating species trees that bypasses MCMC integration over gene trees with sophisticated methods for estimating marginal likelihoods, needed for Bayesian model selection, to provide a rigorous and computationally tractable technique for genome-wide species delimitation. We provide a critical yet simple correction that brings the likelihoods of different species trees, and more importantly their corresponding marginal likelihoods, to the same common denominator, which enables direct and accurate comparisons of competing species delimitation models using Bayes factors. We test this approach, which we call Bayes factor delimitation (*with genomic data; BFD*), using common species delimitation scenarios with computer simulations. Varying the numbers of loci and the number of samples suggest that the approach can distinguish the true model even with few loci and limited samples per species. Misspecification of the prior for population size{theta} has little impact on support for the true model. We apply the approach to West African forest geckos (Hemidactylus fasciatus complex) using genome-wide SNP data data. This new Bayesian method for species delimitation builds on a growing trend for objective species delimitation methods with explicit model assumptions that are easily tested.

Evolutionary Biology

An optimized CRISPR/Cas toolbox for efficient germline and somatic genome engineering in Drosophila

The type II CRISPR/Cas system has recently emerged as a powerful method to manipulate the genomes of various organisms. Here, we report a novel toolbox for high efficiency genome engineering of Drosophila melanogaster consisting of transgenic Cas9 lines and versatile guide RNA (gRNA) expression plasmids. Systematic evaluation reveals Cas9 lines with ubiquitous or germline restricted patterns of activity. We also demonstrate differential activity of the same gRNA expressed from different U6 snRNA promoters, with the previously untested U6:3 promoter giving the most potent effect. Choosing an appropriate combination of Cas9 and gRNA allows targeting of essential and non-essential genes with transmission rates ranging from 25% - 100%. We also provide evidence that our optimized CRISPR/Cas tools can be used for offset nicking-based mutagenesis and, in combination with oligonucleotide donors, to precisely edit the genome by homologous recombination with efficiencies that do not require the use of visible markers. Lastly, we demonstrate a novel application of CRISPR/Cas-mediated technology in revealing loss-of-function phenotypes in somatic cells following efficient biallelic targeting by Cas9 expressed in a ubiquitous or tissue-restricted manner. In summary, our CRISPR/Cas tools will facilitate the rapid evaluation of mutant phenotypes of specific genes and the precise modification of the genome with single nucleotide precision. Our results also pave the way for high throughput genetic screening with CRISPR/Cas.

Genetics

A genome-wide analysis of Cas9 binding specificity using ChIP-seq and targeted sequence capture

Clustered regularly interspaced short palindromic repeat (CRISPR) RNA-guided nucleases have gathered considerable excitement as a tool for genome engineering. However, questions remain about the specificity of their target site recognition. Most previous studies have examined predicted off-target binding sites that differ from the perfect target site by one to four mismatches, which represent only a subset of genomic regions. Here, we use ChIP-seq to examine genome-wide CRISPR binding specificity at gRNA-specific and gRNA-independent sites. For two guide RNAs targeting the murine Snurf gene promoter, we observed very high binding specificity at the intended target site while off-target binding was observed at 2- to 6-fold lower intensities. We also identified significant gRNA-independent off-target binding. Interestingly, we found that these regions are highly enriched in the PAM site, a sequence required for target site recognition by CRISPR. To determine the relationship between Cas9 binding and endonuclease activity, we used targeted sequence capture as a high-throughput approach to survey a large number of the potential off-target sites identified by ChIP-seq or computational prediction. A high frequency of indels was observed at both target sites and one off-target site, while no cleavage activity could be detected at other ChIP-bound regions. Our data is consistent with recent finding that most interactions between the CRISPR nuclease complex and genomic PAM sites are transient and do not lead to DNA cleavage. The interactions are stabilized by gRNAs with good matches to the target sequence adjacent to the PAM site, resulting in target cleavage activity.

Molecular Biology

Extraordinarily wide genomic impact of a selective sweep associated with the evolution of sex ratio distorter suppression

Symbionts that distort their hosts sex ratio by favouring the production and survival of females are common in arthropods. Their presence produces intense Fisherian selection to return the sex ratio to parity, typified by the rapid spread of host suppressor loci that restore male survival/development. In this study, we investigated the genomic impact of a selective event of this kind in the butterfly Hypolimnas bolina. Through linkage mapping we first identified a genomic region that was necessary for males to survive Wolbachia-induced killing. We then investigated the genomic impact of the rapid spread of suppression that converted the Samoan population of this butterfly from a 100:1 female-biased sex ratio in 2001, to a 1:1 sex ratio by 2006. Models of this process revealed the potential for a chromosome-wide selective sweep. To measure the impact directly, the pattern of genetic variation before and after the episode of selection was compared. Significant changes in allele frequencies were observed over a 25cM region surrounding the suppressor locus, alongside generation of linkage disequilibrium. The presence of novel allelic variants in 2006 suggests that the suppressor was introduced via immigration rather than through de novo mutation. In addition, further sampling in 2010 indicated that many of the introduced variants were lost or had reduced in frequency since 2006. We hypothesise that this loss may have resulted from a period of purifying selection - removing deleterious material that introgressed during the initial sweep. Our observations of the impact of suppression of sex ratio distorting activity reveal an extraordinarily wide genomic imprint, reflecting its status as one of the strongest selective forces in nature.

Evolutionary Biology

Benchmarking undedicated cloud computing providers for analysis of genomic datasets.

A major bottleneck in biological discovery is now emerging at the computational level. Cloud computing offers a dynamic means whereby small and medium-sized laboratories can rapidly adjust their computational capacity. We benchmarked two established cloud computing services, Amazon Web Services Elastic MapReduce (EMR) on Amazon EC2 instances and Google Compute Engine (GCE), using publicly available genomic datasets (E.coli CC102 strain and a Han Chinese male genome) and a standard bioinformatic pipeline on a Hadoop-based platform. Wall-clock time for complete assembly differed by 52.9% (95%CI: 27.5-78.2) for E.coli and 53.5% (95%CI: 34.4-72.6) for human genome, with GCE being more efficient than EMR. The cost of running this experiment on EMR and GCE differed significantly, with the costs on EMR being 257.3% (95%CI: 211.5-303.1) and 173.9% (95%CI: 134.6-213.1) more expensive for E.coli and human assemblies respectively. Thus, GCE was found to outperform EMR both in terms of cost and wall-clock time. Our findings confirm that cloud computing is an efficient and potentially cost-effective alternative for analysis of large genomic datasets. In addition to releasing our cost-effectiveness comparison, we present available ready-to-use scripts for establishing Hadoop instances with Ganglia monitoring on EC2 or GCE.

Bioinformatics

Rate and cost of adaptation in the Drosophila genome

Recent studies have consistently inferred high rates of adaptive molecular evolution between Drosophila species. At the same time, the Drosophila genome evolves under different rates of recombination, which results in partial genetic linkage between alleles at neighboring genomic loci. Here we analyze how linkage correlations affect adaptive evolution. We develop a new inference method for adaptation that takes into account the effect on an allele at a focal site caused by neighboring deleterious alleles (background selection) and by neighboring adaptive substitutions (hitchhiking). Using complete genome sequence data and fine-scale recombination maps, we infer a highly heterogeneous scenario of adaptation in Drosophila. In high-recombining regions, about 50% of all amino acid substitutions are adaptive, together with about 20% of all substitutions in proximal intergenic regions. In low-recombining regions, only a small fraction of the amino acid substitutions are adaptive, while hitchhiking accounts for the majority of these changes. Hitchhiking of deleterious alleles generates a substantial collateral cost of adaptation, leading to a fitness decline of about 30/2N per gene and per million years in the lowest-recombining regions. Our results show how recombination shapes rate and efficacy of the adaptive dynamics in eukaryotic genomes.

Evolutionary Biology

GC-content evolution in bacterial genomes: the biased gene conversion hypothesis expands.

The characterization of functional elements in genomes relies on the identification of the footprints of natural selection. In this quest, taking into account neutral evolutionary processes such as mutation and genetic drift is crucial because these forces can generate patterns that may obscure or mimic signatures of selection. In mammals, and probably in many eukaryotes, another such confounding factor called GC-Biased Gene Conversion (gBGC) has been documented. This mechanism generates patterns identical to what is expected under selection for higher GC-content, specifically in highly recombining genomic regions. Recent results have suggested that a mysterious selective force favouring higher GC-content exists in Bacteria but the possibility that it could be gBGC has been excluded. Here, we show that gBGC is probably at work in most if not all bacterial species. First we find a consistent positive relationship between the GC-content of a gene and evidence of intra-genic recombination throughout a broad spectrum of bacterial clades. Second, we show that the evolutionary force responsible for this pattern is acting independently from selection on codon usage, and could potentially interfere with selection in favor of optimal AU-ending codons. A comparison with data from human populations shows that the intensity of gBGC in Bacteria is comparable to what has been reported in mammals. We propose that gBGC is not restricted to sexual Eukaryotes but also widespread among Bacteria and could therefore be an ancestral feature of cellular organisms. We argue that if gBGC occurs in bacteria, it can account for previously unexplained observations, such as the apparent non-equilibrium of base substitution patterns and the heterogeneity of gene composition within bacterial genomes. Because gBGC produces patterns similar to positive selection, it is essential to take this process into account when studying the evolutionary forces at work in bacterial genomes.

Evolutionary Biology