Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,459 records · Page 81Linked to original sources

Early modern human dispersal from Africa: genomic evidence for multiple waves of migration

BackgroundAnthropological and genetic data agree in indicating the African continent as the main place of origin for modern human. However, it is unclear whether early modern humans left Africa through a single, major process, dispersing simultaneously over Asia and Europe, or in two main waves, first through the Arab peninsula into Southern Asia and Oceania, and later through a Northern route crossing the Levant.\n\nResultsHere we show that accurate genomic estimates of the divergence times between European and African populations are more recent than those between Australo-Melanesia and Africa, and incompatible with the effects of a single dispersal. This difference cannot possibly be accounted for by the effects of hybridization with archaic human forms in Australo-Melanesia. Furthermore, in several populations of Asia we found evidence for relatively recent genetic admixture events, which could have obscured the signatures of the earliest processes.\n\nConclusionsWe conclude that the hypothesis of a single major human dispersal from Africa appears hardly compatible with the observed historical and geographical patterns of genome diversity, and that Australo-Melanesian populations seem still to retain a genomic signature of a more ancient divergence from Africa

Evolutionary Biology

Genome-Wide Scan for Adaptive Divergence and Association with Population-Specific Covariates

In population genomics studies, accounting for the neutral covariance structure across population allele frequencies is critical to improve the robustness of genome-wide scan approaches. Elaborating on the BO_SCPLOWAYC_SCPLOWEO_SCPLOWNVC_SCPLOW model, this study investigates several modeling extensions i) to improve the estimation accuracy of the population covariance matrix and all the related measures; ii) to identify significantly overly differentiated SNPs based on a calibration procedure of the XtX statistics; and iii) to consider alternative covariate models for analyses of association with population-specific covariables. In particular, the auxiliary variable model allows to deal with multiple testing issues and, providing the relative marker positions are available, to capture some Linkage Disequilibrium information. A comprehensive simulation study was carried out to evaluate the performances of these different models. Also, when compared in terms of power, robustness and computational efficiency, to five other state-of-the-art genome scan methods (BO_SCPLOWAYC_SCPLOWEO_SCPLOWNVC_SCPLOW2, BO_SCPLOWAYC_SCPLOWSO_SCPLOWCC_SCPLOWEO_SCPLOWNVC_SCPLOW, BO_SCPLOWAYC_SCPLOWSO_SCPLOWCANC_SCPLOW, FLK and LFMM) the proposed approaches proved highly effective. For illustration purpose, genotyping data on 18 French cattle breeds were analyzed leading to the identification of thirteen strong signatures of selection. Among these, four (surrounding the KITLG, KIT, EDN3 and ALB genes) contained SNPs strongly associated with the piebald coloration pattern while a fifth (surrounding PLAG1) could be associated to morphological differences across the populations. Finally, analysis of Pool-Seq data from 12 populations of Littorina saxatilis living in two different ecotypes illustrates how the proposed framework might help addressing relevant ecological issue in non-model species. Overall, the proposed methods define a robust Bayesian framework to characterize adaptive genetic differentiation across populations. The BO_SCPLOWAYC_SCPLOWPO_SCPLOWASSC_SCPLOW program implementing the different models is available at http://www1.montpellier.inra.fr/CBGP/software/baypass/.

Genetics

Population genomic scans reveal novel genes underlie convergent flowering time evolution in the introduced range of Arabidopsis thaliana

A long-standing question in evolutionary biology is whether the evolution of convergent phenotypes results from selection on the same heritable genetic components. Using whole genome sequencing and genome scans, we tested whether the evolution of parallel longitudinal flowering time clines in the native and introduced ranges of Arabidopsis thaliana has a similar genetic basis. We found that common variants of large effect on flowering time in the native range do not appear to have been under recent strong selection in the introduced range. Genes in regions of the genome that are under selection for flowering time are also not enriched for functions related to development or environmental sensing. We instead identified a set of 53 new candidate genes putatively linked to the evolution of flowering time in the species introduced range. A high degree of conditional neutrality of flowering time variants between the native and introduced range may preclude parallel evolution at the level of genes. Overall, neither gene pleiotropy nor available standing genetic variation appears to have restricted the evolution of flowering time in the introduced range to high frequency variants from the native range or to known flowering time pathway genes.

Evolutionary Biology

Genome wide estimates of mutation rates and spectrum in Schizosaccharomyces pombe indicate CpG sites are highly mutagenic despite the absence of DNA methylation

We accumulated mutations for 1952 generations in 79 initially identical, haploid lines of the fission yeast Schizosaccharomyces pombe and then performed whole-genome sequencing to determine the mutation rates and spectrum. We captured 696 spontaneous mutations across the 79 mutation accumulation lines. We compared the mutation spectrum and rate to another model ascomycetous yeast, the budding yeast Saccharomyces cerevisiae. While the two organisms are approximately 600 million years diverged from each other, they share similar life histories, genome size and genomic G/C content. We found that Sc. pombe and S. cerevisiae have similar mutation rates, contrary to what was expected given Sc. pombes smaller reported effective population size. Sc. pombes also exhibits a strong insertion bias in comparison to S. cerevisiae. Intriguingly, we observed an increased mutation rate at cytosine nucleotides, specifically CpG nucleotides, which is also seen in S. cerevisiae. However, the absence of methylation in Sc. pombe and the pattern of mutation at these sites, primarily C[->] A as opposed to C[->]T, strongly suggest that the increased mutation rate is not caused by deamination of methylated cytosines. This result implies that the high mutability of CpG dinucleotides in other species may be caused in part by an additional mechanism than methylation.

Genetics

ProtAnnot: an App for Integrated Genome Browser to display how alternative splicing and transcription affect proteins

SummaryOne gene can produce multiple transcript variants encoding proteins with different functions. To facilitate visual analysis of transcript variants, we developed ProtAnnot, which shows protein annotations in the context of genomic sequence. ProtAnnot searches InterPro and displays profile matches (protein annotations) alongside gene models, exposing how alternative promoters, splicing, and 3 end processing add, remove, or remodel functional motifs. To draw attention to these effects, ProtAnnot color-codes exons by frame and displays a cityscape graphic summarizing exonic sequence at each position. These techniques make visual analysis of alternative transcripts faster and more convenient for biologists. Availability and ImplementationProtAnnot is a plug-in App for Integrated Genome Browser, an open source desktop genome browser available from http://www.bioviz.org. Contactaloraine@uncc.edu

Bioinformatics

Evolutionary dynamics of roX lncRNA function and genomic occupancy

Many long noncoding RNAs (lncRNAs) can regulate chromatin states, but the evolutionary origin and dynamics driving lncRNA-genome interactions are unclear. We developed an integrative strategy that identifies lncRNA orthologs in different species despite limited sequence similarity that is applicable to fly and mammalian lncRNAs. Analysis of the roX lncRNAs, which are essential for dosage compensation of the single X-chromosome in Drosophila males, revealed 47 new roX orthologs in diverse Drosophilid species across ~40 million years of evolution. Genetic rescue by roX orthologs and engineered synthetic lncRNAs showed that evolutionary maintenance of focal structural repeats mediates roX function. Genomic occupancy maps of roX RNAs in four species revealed rapid turnover of individual binding sites but conservation within nearby chromosomal neighborhoods. Many new roX binding sites evolved from DNA encoding a pre-existing RNA splicing signal, effectively linking dosage compensation to transcribed genes. Thus, evolutionary analysis illuminates the principles for the birth and death of lncRNAs and their genomic targets.

Evolutionary Biology

Decomposing variability in protein levels from noisy expression, genome duplication and partitioning errors during cell-divisions

Inside individual cells, expression of genes is inherently stochastic and manifests as cell-to-cell variability or noise in protein copy numbers. Since proteins half-lives can be comparable to the cell-cycle length, randomness in cell-division times generates additional intercellular variability in protein levels. Moreover, as many mRNA/protein species are expressed at low-copy numbers, errors incurred in partitioning of molecules between the mother and daughter cells are significant. We derive analytical formulas for the total noise in protein levels for a general class of cell-division time and partitioning error distributions. Using a novel hybrid approach the total noise is decomposed into components arising from i) stochastic expression; ii) partitioning errors at the time of cell-division and iii) random cell-division events. These formulas reveal that random cell-division times not only generate additional extrinsic noise but also critically affect the mean protein copy numbers and intrinsic noise components. Counter intuitively, in some parameter regimes noise in protein levels can decrease as cell-division times become more stochastic. Computations are extended to consider genome duplication, where the gene dosage is increased by two-fold at a random point in the cell-cycle. We systematically investigate how the timing of genome duplication influences different protein noise components. Intriguingly, results show that noise contribution from stochastic expression is minimized at an optimal genome duplication time. Our theoretical results motivate new experimental methods for decomposing protein noise levels from single-cell expression data. Characterizing the contributions of individual noise mechanisms will lead to precise estimates of gene expression parameters and techniques for altering stochasticity to change phenotype of individual cells.

Systems Biology

Haplotype synthesis analysis in public reference data reveals functional variants underlying known genome-wide associated susceptibility loci

The functional mechanisms underlying disease association identified by Genome-wide Association Studies remain unknown for susceptibility loci located outside gene coding regions. In addition to the regulation of gene expression, synthesis of effects from multiple surrounding functional variants has been suggested as an explanation of hard-to-interpret associations.\n\nHere, we define filter criteria based on linkage disequilibrium measures and allele frequencies which reflect expected properties of synthesizing variant sets. For eligible candidate sets we search for those haplotypes that are highly correlated with the risk alleles of a genome-wide associated variant.\n\nWe applied our methods to 1,000 Genomes reference data and confirmed Crohns Disease and Type 2 Diabetes susceptibility loci. Of these, a proportion of 32% allowed explanation by three-variant-haplotypes carrying at least two functional variants, as compared to a proportion of 16% for random variants (P = 2.92 {middle dot} 10-6). More importantly, we detected examples of known loci whose association can fully be explained by surrounding missense variants: three missense variants from MUC19 synthesize rs11564258 (L0C105369736/MUC19, intron; Crohns Disease). Next, rs2797685 (PER3, intron; Crohns Disease) is synthesized by a 57 kilobase haplotype defined by five missense variants from PER3 and three missense variants from UTS2. Finally, the association of rs7178572 (HMG20A, intron; Type 2 Diabetes) can be explained by the synthesis of eight haplotypes, each carrying at least one missense variant in either PEAK1, TBC1D2B, CHRNA5 or ADAMTS7.\n\nIn summary, application of our new methods highlights the potential of synthesis analysis to guide functional follow-up investigation of findings from association studies.

Bioinformatics

FINEMAP: Efficient variable selection using summary data from genome-wide association studies

MotivationThe goal of fine-mapping in genomic regions associated with complex diseases and traits is to identify causal variants that point to molecular mechanisms behind the associations. Recent fine-mapping methods using summary data from genome-wide association studies rely on exhaustive search through all possible causal configurations, which is computationally expensive.\n\nResultsWe introduce FINEMAP, a software package to efficiently explore a set of the most important causal configurations of the region via a shotgun stochastic search algorithm. We show that FINEMAP produces accurate results in a fraction of processing time of existing approaches and is therefore a promising tool for analyzing growing amounts of data produced in genome-wide association studies.\n\nAvailabilityFINEMAP v1.0 is freely available for Mac OS X and Linux at http://www.christianbenner.com.\n\nContact: christian.benner@helsinki.fi, matti.pirinen@helsinki.fi

Genetics

One Codex: A Sensitive and Accurate Data Platform for Genomic Microbial Identification

High-throughput sequencing (HTS) is increasingly being used for broad applications of microbial characterization, such as microbial ecology, clinical diagnosis, and outbreak epidemiology. However, the analytical task of comparing short sequence reads against the known diversity of microbial life has proved to be computationally challenging. The One Codex data platform was created with the dual goals of analyzing microbial data against the largest possible collection of microbial reference genomes, as well as presenting those results in a format that is consumable by applied end-users. One Codex identifies microbial sequences using a \"k-mer based\" taxonomic classification algorithm through a web-based data platform, using a reference database that currently includes approximately 40,000 bacterial, viral, fungal, and protozoan genomes. In order to evaluate whether this classification method and associated database provided quantitatively different performance for microbial identification, we created a large and diverse evaluation dataset containing 50 million reads from 10,639 genomes, as well as sequences from six organisms novel species not be included in the reference databases of any of the tested classifiers. Quantitative evaluation of several published microbial detection methods shows that One Codex has the highest degree of sensitivity and specificity (AUC = 0.97, compared to 0.82-0.88 for other methods), both when detecting well-characterized species as well as newly sequenced, \"taxonomically novel\" organisms.

Bioinformatics

Comparing cancer cell lines and tumor samples by genomic profiles

Cancer cell lines are often used in laboratory experiments as models of tumors, although they can have substantially different genetic and epigenetic profiles compared to tumors. We have developed a general computational method - TumorComparer - to systematically quantify similarities and differences between tumor material when detailed genetic and molecular profiles are available. The comparisons can be flexibly tailored to a particular biological question by placing a higher weight on functional alterations of interest ( weighted similarity). In a first pan-cancer application, we have compared 260 cell lines from the Cancer Cell Line Encyclopaedia (CCLE) and 1914 tumors of six different cancer types from The Cancer Genome Atlas (TCGA), using weights to emphasize genomic alterations that frequently recur in tumors. We report the potential suitability of particular cell lines as tumor models and identify apparently unsuitable outlier cell lines, some of which are in wide use, for each of the six cancer types. In future, this weighted similarity method may be generalized for use in a clinical setting to compare patient profiles consisting of genomic patterns combined with clinical attributes, such as diagnosis, treatment and response to therapy.

Cancer Biology

The Great Migration and African-American genomic diversity

Genetic studies of African-Americans identify functional variants, elucidate historical and genealogical mysteries, and reveal basic biology. However, African-Americans have been under-represented in genetic studies, and little is known about nation-wide patterns of genomic diversity in the population. Here, we present a comprehensive assessment of African-American genomic diversity using genotype data from nationally and regionally representative cohorts. We find higher African ancestry in southern United States compared to the North and West. We show that relatedness patterns track north- and west-bound routes followed during the Great Migration, suggesting that admixture occurred predominantly in the South prior to the Civil War and that ancestry-biased migration is responsible for regional differences in ancestry. Rare genetic traits among African-Americans can therefore be shared over long geographic distances along the Great Migration routes, yet their distribution over short distances remains highly structured. This study clarifies the role of recent demography in shaping African-American genomic diversity.

Preprint

Whole-genome modeling accurately predicts quantitative traits, as revealed in plants.

Many adaptive events in natural populations, as well as response to artificial selection, are caused by polygenic action. Under selective pressure, the adaptive traits can quickly respond via small allele frequency shifts spread across numerous loci. We hypothesize that a large proportion of current phenotypic variation between individuals may be best explained by population admixture.\n\nWe thus consider the complete, genome-wide universe of genetic variability, spread across several ancestral populations originally separated. We experimentally confirmed this hypothesis by predicting the differences in quantitative disease resistance levels among accessions in the wild legume Medicago truncatula. We discovered also that variation in genome admixture proportion explains most of phenotypic variation for several quantitative functional traits, but not for symbiotic nitrogen fixation. We shown that positive selection at the species level might not explain current, rapid adaptation.\n\nThese findings prove the infinitesimal model as a mechanism for adaptation of quantitative phenotypes. Our study produced the first evidence that the whole-genome modeling of DNA variants is the best approach to describe an inherited quantitative trait in a higher eukaryote organism and proved the high potential of admixture-based analyses. This insight contribute to the understanding of polygenic adaptation, and can accelerate plant and animal breeding, and biomedicine research programs.

Genetics

Whole genome duplication in coast redwood (Sequoia sempervirens) and its implications for explaining the rarity of polyploidy in conifers

SO_SCPLOWUMMARYC_SCPLOWO_LIWhereas polyploidy is common and an important evolutionary factor in most land plant lineages it is a real rarity in gymnosperms. Coast redwood (Sequoia sempervirens) is the only hexaploid conifer and one of just two naturally polyploid conifer species. Numerous hypotheses about the mechanism of polyploidy in Sequoia and parental genome donors have been proffered over the years, primarily based on morphological and cytological data, but it remains unclear how Sequoia became polyploid and why this lineage overcame an apparent gymnosperm barrier to whole-genome duplication (WGD).\nC_LIO_LIWe sequenced transcriptomes and used phylogenetic inference, Bayesian concordance analysis, and paralog age distributions to resolve relationships among gene copies in hexaploid coast redwood and its close relatives.\nC_LIO_LIOur data show that hexaploidy in the coast redwood lineage is best explained by autopolyploidy or, if there was allopolyploidy, this was restricted to within the Californian redwood clade. We found that duplicate genes have more similar sequences than would be expected given evidence from fossil guard cell size which suggest that polyploidy dates to the Eocene.\nC_LIO_LIConflict between molecular and fossil estimates of WGD can be explained if diploidization occurred very slowly following whole genome duplication. We extrapolate from this to suggest that the rarity of polyploidy in conifers may be due to slow rates of diploidization in this clade.\nC_LI

Plant Biology

CRISPResso: sequencing analysis toolbox for CRISPR genome editing

To the Editor To the Editor References Recent progress in genome editing technologies, in particular the CRISPR-Cas9 nuclease system, has provided new opportunities to investigate the biological functions of genomic sequences by targeted mutagenesis [1-4]. Briefly, Cas9 may be directed by a chimeric single guide RNA (sgRNA) to a target genomic sequence upstream of a protospacer adjacent motif (PAM) for cleavage. Double strand breaks (DSBs) resulting from site-specific Cas9 cleavage can be resolved by endogenous DNA repair pathways such as non-homologous end joining (NHEJ) or homology-directed repair (HDR). These repair mechanisms result in a spectrum of diverse outcomes including insertions, deletions, nucleotide substitutions, and, in the case of HDR, recombination of extrachromosomal donor sequences [1-3, 5 ...

Bioinformatics

AGOUTI: improving genome assembly and annotation using transcriptome data

SummaryCurrent genome assemblies consist of thousands of contigs. These incomplete and fragmented assemblies lead to errors in gene identification, such that single genes spread across multiple contigs are annotated as separate gene models. We present AGOUTI (Annotated Genome Optimization Using Transcriptome Information), a tool that uses RNA-seq data to simultaneously combine contigs into scaffolds and fragmented gene models into single models. We show that AGOUTI improves both the contiguity of genome assemblies and the accuracy of gene annotation, providing updated versions of each as output.\n\nAvailabilityThe software is implemented in python and is available from github.com/svm-zhang/AGOUTI.\n\nContactsimozhan@indiana.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

Bioinformatics

Interactive Analytics for Very Large Scale Genomic Data

Large scale genomic sequencing is now widely used to decipher questions in diverse realms such as biological function, human diseases, evolution, ecosystems, and agriculture. With the quantity and diversity these data harbor, a robust and scalable data handling and analysis solution is desired. Here we present interactive analytics using public cloud infrastructure and distributed computing database Dremel and developed according to the standards of Global Alliance for Genomics and Health, to perform information compression, comprehensive quality controls, and biological information retrieval in large volumes of genomic data. We demonstrate that such computing paradigms can provide orders of magnitude faster turnaround for common analyses, transforming long-running batch jobs submitted via a Linux shell into questions that can be asked from a web browser in seconds.

Preprint

Comparative evaluation of the genomes of common bacterial members of the Drosophila intestinal community

Drosophila melanogaster is an excellent model to explore the molecular exchanges that occur between an animal intestine and their microbial passengers. For example, groundbreaking studies in flies uncovered a sophisticated web of host responses to intestinal bacteria. The outcomes of these responses define critical events in the host, such as the establishment of immune responses, access to nutrients, and the rate of larval development. Despite our steady march towards illuminating the host machinery that responds to bacterial presence in the gut, we know remarkably little about the microbial products that influence bacterial association with a fly host. To address this deficiency, we sequenced and characterized the genomes of three common Drosophila-associated microbes: Lactobacillus plantarum, Lactobacillus brevis and Acetobacter pasteurianus. In each case, we compared the genomes of Drosophila-associated strains to the genomes of strains isolated from alternative sources. This approach allowed us to identify molecular functions common to Drosophila-associated microbes, and, in the case of A. pasteurianus, to identify genes that are essential for association with the host. Of note, many of the gene products unique to fly-associated strains have established roles in the stabilization of host-microbe interactions. We believe that these data provide a valuable starting point for a more thorough examination of the microbial perspective on host-microbe relationships.

Microbiology