Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,027 records · Page 57Linked to original sources

Efficient Breeding by Genomic Mating

In this article, we propose an approach to breeding which focuses on mating instead of truncation selection, our method uses genome-wide marker information in a similar fashion to genomic selection so we refer it to as genomic mating. Using concepts of estimated breeding values, risk (usefulness) and inbreeding, an efficient mating approach is formulated for improvement of breeding values in the long run. We have used a genetic algorithm to find solutions to this optimization problem. Results from our simulations point to the efficiency of genomic mating for breeding complex traits compared to truncation selection.

Genetics

A fast and accurate method for detection of IBD shared haplotypes in genome-wide SNP data

Identical by descent (IBD) segments are used to understand a number of fundamental issues in genetics. IBD segments are typically detected using long stretches of identical alleles between haplotypes in whole-genome SNP data. Phase or SNP call errors in genomic data can degrade accuracy of IBD detection and lead to false positive calls, false negative calls, and under- or overextension of true IBD segments. Furthermore, the number of comparisons increases quadratically with sample size, requiring high computational efficiency. We developed a new IBD segment detection program, FISHR (Find IBD Shared Haplotypes Rapidly), in an attempt to accurately detect IBD segments and to better estimate their endpoints using an algorithm that is fast enough to be deployed on the very large whole-genome SNP datasets. We compared the performance of FISHR to three leading IBD segment detection programs: GERMLINE, refinedIBD, and HaploScore. Using simulated and real genomic sequence data, we show that FISHR is slightly more accurate than all programs at detecting long (>3 cM) IBD segments but slightly less accurate than refinedIBD at detecting short (~1 cM) IBD segments. Moreover, FISHR outperforms all programs in determining the true endpoints of IBD segments, which is important for several reasons. FISHR takes two to four times longer than GERMLINE to run, whereas both GERMLINE and FISHR were orders of magnitude faster than refinedIBD and HaploScore. Overall, FISHR provides accurate IBD detection in unrelated individuals and is computationally efficient enough to be utilized on large SNP datasets > 20,000 individuals.

Bioinformatics

Wedding higher taxonomic ranks with metabolic signatures coded in prokaryotic genomes

Taxonomy of prokaryotes has remained a controversial discipline due to the extreme plasticity of microorganisms, causing inconsistencies between phenotypic and genotypic classifications. The genomics era has enhanced taxonomy but also opened new debates about the best practices for incorporating genomic data into polyphasic taxonomy protocols, which are fairly biased towards the identification of bacterial species. Here we use an extensive dataset of Archaea and Bacteria to prove that metabolic signatures coded in their genomes are informative traits that allow to accurately classify organisms coherently to higher taxonomic ranks, and to associate functional features with the definition of taxa. Our results support the ecological coherence of higher taxonomic ranks and reconciles taxonomy with traditional chemotaxonomic traits inferred from genomes. KARL, a simple and free tool useful for assisting polyphasic taxonomy or to perform functional prospections is also presented (https://github.com/giraola/KARL).

Microbiology

Genomic evidence for population-specific responses to coevolving parasites in a New Zealand freshwater snail

Reciprocal coevolving interactions between hosts and parasites are a primary source of strong selection that can promote rapid and often population- or genotype-specific evolutionary change. These host-parasite interactions are also a major source of disease. Despite their importance, very little is known about the genomic basis of coevolving host-parasite interactions in natural populations, especially in animals. Here, we use gene expression and sequence evolution approaches to take critical steps towards characterizing the genomic basis of interactions between the freshwater snail Potamopyrgus antipodarum and its coevolving sterilizing trematode parasite, Microphallus sp., a textbook example of natural coevolution. We found that Microphallus-infected P. antipodarum exhibit systematic downregulation of genes relative to uninfected P. antipodarum. The specific genes involved in parasite response differ markedly across lakes, consistent with a scenario where population-level coevolution is leading to population-specific host-parasite interactions and evolutionary trajectories. We also used an FST-based approach to identify a set of loci that represent promising candidates for targets of parasite-mediated selection across lakes as well as within each lake population. These results constitute the first genomic evidence for population-specific responses to coevolving infection in the P. antipodarum-Microphallus interaction and provide new insights into the genomic basis of coevolutionary interactions in nature.

Evolutionary Biology

CNView: a visualization and annotation tool for copy number variation from whole-genome sequencing

SummaryCopy number variation (CNV) is a major component of structural differences between individual genomes. The recent emergence of population-scale whole-genome sequencing (WGS) datasets has enabled genome-wide CNV delineation. However, molecular validation at this scale is impractical, so visualization is an invaluable preliminary screening approach when evaluating CNVs. Standardized tools for visualization of CNVs in large WGS datasets are therefore in wide demand.\n\nMethods & ResultsTo address this demand, we developed a software tool, CNView, for normalized visualization, statistical scoring, and annotation of CNVs from population-scale WGS datasets. CNView surmounts challenges of sequencing depth variability between individual libraries by locally adapting to cohort-wide variance in sequencing uniformity at any locus. Importantly, CNView is broadly extensible to any reference genome assembly and most current WGS data types.\n\nAvailability and ImplementationCNView is written in R, is supported on OS X, MS Windows, and Linux, and is freely distributed under the MIT license. Source code and documentation are available from https://github.com/RCollins13/CNView\n\nContacttalkowski@chgr.mgh.harvard.edu

Bioinformatics

Whole-genome sequencing of an advanced case of small-cell gallbladder neuroendocrine carcinoma

The majority of gallbladder cancer cases are discovered at later stages, which frequently leads to poor prognoses. Small-cell gallbladder neuroendocrine carcinoma (GB-SCNEC) is a relatively rare histological type of gallbladder cancer, and its survival rate is exceptionally low because of its greater malignant potential. In addition, the genomic landscape of GB-SCNEC is rarely considered in treatment decisions. We performed whole-genome sequencing on an advanced case of GB-SCNEC. By analyzing the whole-genome sequencing data of the primary cancer tissue (76.29X coverage), lymphatic metastatic cancer tissue (73.92X coverage) and matched non-cancerous tissue (35.73X coverage), we identified approximately 900 high-quality somatic single nucleotide variants (SNVs), 109 of which were shared by both the primary and metastatic tumor tissues. Somatic non-synonymous coding variations with damaging impact in HMCN1 and CDH10 were observed in both the primary and metastatic tissue specimens. A pathway analysis of the genes mapped to the SNVs revealed gene enrichment associated with axon guidance, ERBB signaling, sulfur metabolism and calcium signaling. Furthermore, we identified 20 chromosomal rearrangements that included 11 deletions, 4 tandem duplications and 5 inversions that mapped to known genes. Two gene fusions, NCAM2-SGCZ and BTG3-CCDC40 were also discovered and validated by Sanger sequencing. Additionally, we identified genome-wide copy number variations and microsatellite instability. In this study, we identified novel biological markers of GB-SCNEC that may serve as valuable prognostic factors or indicators of treatment response in patients with GB-SCNEC with lymphatic metastasis.

Genetics

pcadapt: an R package to perform genome scans for selection based on principal component analysis

The R package pcadapt performs genome scans to detect genes under selection based on population genomic data. It assumes that candidate markers are outliers with respect to how they are related to population structure. Because population structure is ascertained with principal component analysis, the package is fast and works with large-scale data. It can handle missing data and pooled sequencing data. By contrast to population-based approaches, the package handle admixed individuals and does not require grouping individuals into populations. Since its first release, pcadapt has evolved both in terms of statistical approach and software implementation. We present results obtained with robust Mahalanobis distance, which is a new statistic for genome scans available in the 2.0 and later versions of the package. When hierarchical population structure occurs, Mahalanobis distance is more powerful than the communality statistic that was implemented in the first version of the package. Using simulated data, we compare pcadapt to other software for genome scans (BayeScan, hapflk, OutFLANK, sNMF). We find that the proportion of false discoveries is around a nominal false discovery rate set at 10% with the exception of BayeScan that generates 40% of false discoveries. We also find that the power of BayeScan is severely impacted by the presence of admixed individuals whereas pcadapt is not impacted. Last, we find that pcadapt and hapflk are the most powerful software in scenarios of population divergence and range expansion. Because pcadapt handles next-generation sequencing data, it is a valuable tool for data analysis in molecular ecology.

Ecology

Linking comparative genomics and environmental distribution patterns of microbial populations through metagenomics

Combining well-established practices from comparative genomics and the emerging opportunities from assembly-based metagenomics can enhance the utility of increasing number of metagenome-assembled genomes (MAGs). Here we used protein clustering to characterize 48 MAGs and 10 cultivars based on their entire gene content, and linked this information to their environmental distribution patterns to better understand the microbial response to the 2010 Deepwater Horizon oil spill in the Gulf of Mexico coastline. Our results suggest that while most oil-associated bacterial populations originated from the ocean, a few actually emerged from the sand rare biosphere. These new findings suggest that there are considerable benefits to employ approaches from comparative genomics to study the whole content of newly identified genomes, and the investigation of emerging patterns in the environmental context can augment the efficacy of assembly-based metagenomic surveys.

Microbiology

Evolutionary genomics of peach and almond domestication

The domesticated almond [Prunus dulcis (L.) Batsch] and peach [P. persica (Mill.) D. A. Webb] originate on opposite sides of Asia and were independently domesticated approximately 5000 years ago. While interfertile, they possess alternate mating systems and differ in a number of morpholog-ical and physiological traits. Here we evaluated patterns of genome-wide diversity in both almond and peach to better understand the impacts of mating system, adaptation, and domestication on the evolution of these taxa. Almond has [~]7X the genetic diversity of peach, and high genome-wide FST values support their status as separate species. We estimated a divergence time of approximately 8 Mya, coinciding with an active period of uplift in the northeast Tibetan Plateau and subsequent Asian climate change. We see no evidence of bottleneck during domestication of either species, but identify a number of regions showing signatures of selection during domestication and a significant overlap in candidate regions between peach and almond. While we expected gene expression in fruit to overlap with candidate selected regions, instead we find enrichment for loci highly differentiated between the species, consistent with recent fossil evidence suggesting fruit divergence long preceded domestication. Taken together this study tells us how closely related tree species evolve and are domesticated, the impact of these events on their genomes, and the utility of genomic information for long-lived species. Further exploration of this data will contribute to the genetic knowledge of these species and provide information regarding targets of selection for breeding application and further the understanding of evolution in these species.

Evolutionary Biology

Hybrid assembly of the large and highly repetitive genome of Aegilops tauschii, a progenitor of bread wheat, with the mega-reads algorithm

Long sequencing reads generated by single-molecule sequencing technology offer the possibility of dramatically improving the contiguity of genome assemblies. The biggest challenge today is that long reads have relatively high error rates, currently around 15%. The high error rates make it difficult to use this data alone, particularly with highly repetitive plant genomes. Errors in the raw data can lead to insertion or deletion errors (indels) in the consensus genome sequence, which in turn create significant problems for downstream analysis; for example, a single indel may shift the reading frame and incorrectly truncate a protein sequence. Here we describe an algorithm that solves the high error rate problem by combining long, high-error reads with shorter but much more accurate Illumina sequencing reads, whose error rates average <1%. Our hybrid assembly algorithm combines these two types of reads to construct mega-reads, which are both long and accurate, and then assembles the mega-reads using the CABOG assembler, which was designed for long reads. We apply this technique to a large data set of Illumina and PacBio sequences from the species Aegilops tauschii, a large and highly repetitive plant genome that has resisted previous attempts at assembly. We show that the resulting assembled contigs are far larger than in any previous assembly, with an N50 contig size of 486,807. We compare the contigs to independently produced optical maps to evaluate their large-scale accuracy, and to a set of high-quality bacterial artificial chromosome (BAC)-based assemblies to evaluate base-level accuracy.

Bioinformatics

Genome-wide biases in the rate and molecular spectrum of spontaneous mutations in Vibrio cholerae and Vibrio fischeri

The vast diversity in nucleotide composition and architecture among bacterial genomes may be partly explained by inherent biases in the rates and spectra of spontaneous mutations. Bacterial genomes with multiple chromosomes are relatively unusual but some are relevant to human health, none more so than the causative agent of cholera, Vibrio cholerae. Here, we present the genome-wide mutation spectra in wild-type and mismatch repair (MMR) defective backgrounds of two Vibrio species, the low-GC% squid symbiont V. fischeri and the pathogen V. cholerae, collected under conditions that greatly minimize the efficiency of natural selection. In apparent contrast to their high diversity in nature, both wild-type V. fischeri and V. cholerae have among the lowest rates for base-substitution mutations (bpsms) and insertion-deletion mutations (indels) that have been measured, below 10-3/genome/generation. V. fischeri and V. cholerae have distinct mutation spectra, but both are AT-biased and produce a surprising number of multi-nucleotide indels. Furthermore, the loss of a functional MMR system caused the mutation spectra of these species to converge, implying that the MMR system itself contributes to species-specific mutation patterns. Bpsm and indel rates varied among genome regions, but do not explain the more rapid evolutionary rates of genes on chromosome 2, which likely result from weaker purifying selection. More generally, the very low mutation rates of Vibrio species correlate inversely with their immense population sizes and suggest that selection may not only have maximized replication fidelity but also optimized other polygenic traits relative to the constraints of genetic drift.

Genetics

Extensive genetic diversity among populations of the malaria mosquito Anopheles moucheti revealed by population genomics

Malaria vectors are exposed to intense selective pressures due to large-scale intervention programs that are underway in most African countries. One of the current priorities is therefore to clearly assess the adaptive potential of Anopheline populations, which is critical to understand and anticipate the response mosquitoes can elicit against such adaptive challenges. The development of genomic resources that will empower robust examinations of evolutionary changes in all vectors including currently understudied species is an inevitable step toward this goal. Here we constructed double-digest Restriction Associated DNA (ddRAD) libraries and generated 6461 Single Nucleotide Polymorphisms (SNPs) that we used to explore the population structure and demographic history of wild-caught Anopheles moucheti from Cameroon. The genome-wide distribution of allelic frequencies among samples best fitted that of an old population at equilibrium, characterized by a weak genetic structure and extensive genetic diversity, presumably due to a large long term effective population size. Estimates of FST and Linkage Disequilibrium (LD) across SNPs reveal a very low genetic differentiation throughout the genome and the absence of segregating LD blocks among populations, suggesting an overall lack of local adaptation. Our study provides the first investigation of the genetic structure and diversity in An. moucheti at the genomic scale. We conclude that, despite a weak genetic structure, this species has the potential to challenge current vector control measures and other rapid anthropogenic and environmental changes thanks to its great genetic diversity.

Evolutionary Biology

Genome-scale rates of evolutionary change in bacteria

Estimating the rates at which bacterial genomes evolve is critical to understanding major evolutionary and ecological processes such as disease emergence, long-term host-pathogen associations, and short-term transmission patterns. The surge in bacterial genomic data sets provides a new opportunity to estimate these rates and reveal the factors that shape bacterial evolutionary dynamics. For many organisms estimates of evolutionary rate display an inverse association with the time-scale over which the data are sampled. However, this relationship remains unexplored in bacteria due to the difficulty in estimating genome-wide evolutionary rates, which are impacted by the extent of temporal structure in the data and the prevalence of recombination. We collected 36 whole genome sequence data sets from 16 species of bacterial pathogens to systematically estimate and compare their evolutionary rates and assess the extent of temporal structure in the absence of recombination. The majority (28/36) of data sets possessed sufficient clock-like structure to robustly estimate evolutionary rates. However, in some species reliable estimates were not possible even with \"ancient DNA\" data sampled over many centuries, suggesting that they evolve very slowly or that they display extensive rate variation among lineages. The robustly estimated evolutionary rates spanned several orders of magnitude, from 10-6 to 10-8 nucleotide substitutions site-1 year-1. This variation was largely attributable to sampling time, which was strongly negatively associated with estimated evolutionary rates, with this relationship best described by an exponential decay curve. To avoid potential estimation biases such time-dependency should be considered when inferring evolutionary time-scales in bacteria.

Evolutionary Biology

Reconstructing the backbone of the Saccharomycotina yeast phylogeny using genome-scale data

Understanding the phylogenetic relationships among the yeasts of the subphylum Saccharomycotina is a prerequisite for understanding the evolution of their metabolisms and ecological lifestyles. In the last two decades, the use of rDNA and multi-locus data sets has greatly advanced our understanding of the yeast phylogeny, but many deep relationships remain unsupported. In contrast, phylogenomic analyses have involved relatively few taxa and lineages that were often selected with limited considerations for covering the breadth of yeast biodiversity. Here we used genome sequence data from 86 publicly available yeast genomes representing 9 of the 11 major lineages and 10 non-yeast fungal outgroups to generate a 1,233-gene, 96-taxon data matrix. Species phylogenies reconstructed using two different methods (concatenation and coalescence) and two data matrices (amino acids or the first two codon positions) yielded identical and highly supported relationships between the 9 major lineages. Aside from the lineage comprised by the family Pichiaceae, all other lineages were monophyletic. Most interrelationships among yeast species were robust across the two methods and data matrices. However, 8 of the 93 internodes conflicted between analyses or data sets, including the placements of: the clade defined by species that have reassigned the CUG codon to encode serine, instead of leucine; the clade defined by a whole genome duplication; and of Ascoidea rubescens. These phylogenomic analyses provide a robust roadmap for future comparative work across the yeast subphylum in the disciplines of taxonomy, molecular genetics, evolutionary biology, ecology, and biotechnology. To further this end, we have also provided a BLAST server to query the 86 Saccharomycotina genomes, which can be found at http://y1000plus.org/blast.

Evolutionary Biology

Assembly of Radically Recoded E. coli Genome Segments

The large potential of radically recoded organisms (RROs) in medicine and industry depends on improved technologies for efficient assembly and testing of recoded genomes for biosafety and functionality. Here we describe a next generation platform for conjugative assembly genome engineering, termed CAGE 2.0, that enables the scarless integration of large synthetically recoded E. coli segments at isogenic and adjacent genomic loci. A stable tdk dual selective marker is employed to facilitate cyclical assembly and removal of attachment sites used for targeted segment delivery by sitespecific recombination. Bypassing the need for vector transformation harnesses the multi Mb capacity of CAGE, while minimizing artifacts associated with RecA-mediated homologous recombination. Our method expands the genome engineering toolkit for radical modification across many organisms and recombinase-mediated cassette exchange (RMCE).

Synthetic Biology

The impacts of drift and selection on genomic evolution in holometabolous insects

Genomes evolve through a medley of mutation, drift, and selection, all of which act heterogeneously across genes and lineages. The pacemaker models of genomic evolution describe the resulting patterns of evolutionary rate variation: genes that are governed by the same pacemaker exhibit the same pattern of rate heterogeneity across lineages. However, the relative importance of drift and selection in determining the structure of these pacemakers is unknown. Here, we propose a novel phylogenetic approach to explain the formation of pacemakers. We apply this method to a genomic dataset from holometabolous insects, an ancient and diverse group of organisms. We show that when drift is the dominant evolutionary process, each pacemaker tends to govern a large number of fast-evolving genes. In contrast, strong negative selection leads to many distinct pacemakers, each of which governs a few slow-evolving genes. Our results provide new insights into the interplay between drift and selection in driving genomic evolution.

Evolutionary Biology

Genes mirror migrations and cultures in prehistoric Europe - a population genomic perspective

Genomic information from ancient human remains is beginning to show its full potential for learning about human prehistory. We review the last few years' dramatic finds about European prehistory based on genomic data from humans that lived many millennia ago and relate it to modern-day patterns of genomic variation. The early times, the Upper Palaeolithic, appears to contain several population turn-overs followed by more stable populations after the Last Glacial Maximum and during the Mesolithic. Some 11,000 years ago the migrations driving the Neolithic transition start from around Anatolia and reach the north and the west of Europe millennia later followed by major migrations during the Bronze age. These findings show that culture and lifestyle were major determinants of genomic differentiation and similarity in pre-historic Europe rather than geography as is the case today.

Evolutionary Biology

LoRTE: Detecting transposon-induced genomic variants using low coverage PacBio long read sequences

MotivationPopulation genomic analysis of transposable elements has greatly benefited from recent advances of sequencing technologies. However, the propensity of transposable elements to nest in highly repeated regions of genomes limits the efficiency of bioinformatic tools when short read sequences technology is used.\n\nResultsLoRTE is the first tool able to use PacBio long read sequences to identify transposon deletions and insertions between a reference genome and genomes of different strains or populations. Tested against Drosophila melanogaster PacBio datasets, LoRTE appears to be a reliable and broadly applicable tools to study the dynamic and evolutionary impact of transposable elements using low coverage, long read sequences.\n\nAvailability and ImplementationLoRTE is available at http://www.egce.cnrs-gif.fr/?p=6422. It is written in Python 2.7 and only requires the NCBI BLAST + package. LoRTE can be used on standard computer with limited RAM resources and reasonable running time even with large datasets.\n\nContactjonathan.filee@ecge.cnrs-gif.fr

Bioinformatics