Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,279 records · Page 71Linked to original sources

Whole genome sequencing enables definitive diagnosis of Cystic Fibrosis and Primary Ciliary Dyskinesia

Understanding the genomic basis of inherited respiratory disorders can assist in the clinical management of individuals with these rare disorders. We apply whole genome sequencing for the discovery of disease-causing variants in the non-coding regions of known disease genes for two individuals with inherited respiratory disorders. We describe analysis strategies to pinpoint candidate non-coding variants within the non-coding genome and demonstrate aberrant RNA splicing as a result of deep intronic variants in DNAH11 and CFTR. These findings confirm clinical diagnoses of primary ciliary dyskinesia and cystic fibrosis, respectively.

genomics

SRST2: Rapid genomic surveillance for public health and hospital microbiology labs

Rapid molecular typing of bacterial pathogens is critical for public health epidemiology, surveillance and infection control, yet routine use of whole genome sequencing (WGS) for these purposes poses significant challenges. Here we present SRST2, a read mapping-based tool for fast and accurate detection of genes, alleles and multi-locus sequence types (MLST) from WGS data. Using >900 genomes from common pathogens, we show SRST2 is highly accurate and outperforms assembly-based methods in terms of both gene detection and allele assignment. Here we have demonstrated the use of SRST2 for microbial genome surveillance in a variety of public health and hospital settings. In the face of rising threats of antimicrobial resistance and emerging virulence amongst bacterial pathogens, SRST2 represents a powerful tool for rapidly extracting clinically useful information from raw WGS data. Source code is available from http://katholt.github.io/srst2/.

Genomics

Long-read, whole genome shotgun sequence data for five model organisms

Single molecule, real-time (SMRT) sequencing from Pacific Biosciences is increasingly used in many areas of biological research including de novo genome assembly, structural-variant identification, haplotype phasing, mRNA isoform discovery, and base-modification analyses. High-quality, public datasets of SMRT sequences can spur development of analytic tools that can accommodate unique characteristics of SMRT data (long read lengths, lack of GC or amplification bias, and a random error profile leading to high consensus accuracy). In this paper, we describe eight high-coverage SMRT sequence datasets from five organisms (Escherichia coli, Saccharomyces cerevisiae, Neurospora crassa, Arabidopsis thaliana, and Drosophila melanogaster) that have been publicly released to the general scientific community (NCBI Sequence Read Archive ID SRP040522). Data were generated using two sequencing chemistries (P4C2 and P5C3) on the PacBio RS II instrument. The datasets reported here can be used without restriction by the research community to generate whole-genome assemblies, test new algorithms, investigate genome structure and evolution, and identify base modifications in some of the most widely-studied model systems in biological research.

Genomics

msCentipede: Modeling heterogeneity across genomic sites improves accuracy in the inference of transcription factor binding

MotivationUnderstanding global gene regulation depends critically on accurate annotation of regulatory elements that are functional in a given cell type. CENTIPEDE, a powerful, probabilistic framework for identifying transcription factor binding sites from tissue-specific DNase I cleavage patterns and genomic sequence content, leverages the hypersensitivity of factor-bound chromatin and the information in the DNase I spatial cleavage profile characteristic of each DNA binding protein to accurately infer functional factor binding sites. However, the model for the spatial profile in this framework underestimates the substantial variation in the DNase I cleavage profiles across factor-bound genomic locations and across replicate measurements of chromatin accessibility.\n\nResultsIn this work, we adapt a multi-scale modeling framework for inhomogeneous Poisson processes to better model the underlying variation in DNase I cleavage patterns across genomic locations bound by a transcription factor. In addition to modeling variation, we also model spatial structure in the heterogeneity in DNase I cleavage patterns for each factor. Using DNase-seq measurements assayed in a lymphoblastoid cell line, we demonstrate the improved performance of this model for several transcription factors by comparing against the Chip-Seq peaks for those factors. Finally, we propose an extension to this framework that allows for a more flexible background model and evaluate the additional gain in accuracy achieved when the background model parameters are estimated using DNase-seq data from naked DNA. The proposed model can also be applied to paired-end ATAC-seq and DNase-seq data in a straightforward manner.\n\nAvailabilitymsCentipede, a Python implementation of an algorithm to infer transcription factor binding using this model, is made available at https://github.com/rajanil/msCentipede

Genomics

The Avian RNAseq Consortium: a community effort to annotate the chicken genome

Here we describe how members of the chicken research community have come together as the \"Avian RNAseq Consortium\" to provide chicken RNAseq data with a view to improving the annotation of the chicken genome. The data was used by Ensembl in their release 71 gene_build to help provide the most up-to-date annotation of Galgal4, which is still in current use. The data has also been used to predict many ncRNAs, particularly lncRNAs, and it continues to be used as a resource for annotation of other avian genomes along with continued gene discovery in the chicken. This article is submitted as part of the Third Report on Chicken Genes and Chromosomes which will be published in Cytogenetic and Genome Research

Genomics

Genomic prediction of celiac disease targeting HLA-positive individuals

BackgroundGenomic prediction aims to leverage genome-wide genetic data towards better disease diagnostics and risk scores. We have previously published a genomic risk score (GRS) for celiac disease (CD), a common and highly heritable autoimmune disease, which differentiates between CD cases and population-based controls at a clinically-relevant predictive level, improving upon other gene-based approaches. HLA risk haplotypes, particularly HLA-DQ2.5, are necessary but not sufficient for CD, with at least one HLA risk haplotype present in up to half of most Caucasian populations. Here, we assess a genomic prediction strategy that specifically targets this common genetic susceptibility subtype, utilizing a supervised learning procedure for CD that leverages known HLA-DQ2.5 risk.\n\nMethodsUsing L1/L2-regularized support-vector machines trained on large European case-control datasets, we constructed novel CD GRSs specific to individuals with HLA-DQ2.5 risk haplotypes (GRS-DQ2.5) and compared them with the predictive power of the existing CD GRS (GRS14) as well as two haplotype-based approaches, externally validating the results in a North American case-control study.\n\nResultsConsistent with previous observations, both the existing GRS14 and the GRS-DQ2.5 had better predictive performance than the HLA haplotype approaches. GRS-DQ2.5 models, based on directly genotyped or imputed markers, achieved similar levels of predictive performance (AUC = 0.718--0.73), which were substantially higher than those obtained from the DQ2.5 zygosity alone (AUC = 0.558), the HLA risk haplotype method (AUC = 0.634), or the generic GRS14 (AUC = 0.679). In a screening model of at-risk individuals, the GRS-DQ2.5 lowered the number of unnecessary follow-up tests for CD across most sensitivity levels. Relative to a baseline implicating all DQ2.5-positive individuals for follow-up, the GRS-DQ2.5 resulted in a net saving of 2.2 unnecessary follow-up tests for each justified test while still capturing 90% of DQ2.5-positive CD cases.\n\nConclusionsGenomic risk scores for CD that target genetically at-risk sub-groups improve predictive performance beyond traditional approaches and may represent a useful strategy for prioritizing individuals at increase risk of disease, thus potentially reducing unnecessary follow-up diagnostic tests.

Genomics

Distance from Sub-Saharan Africa Predicts Mutational Load in Diverse Human Genomes

The Out-of-Africa (OOA) dispersal ~50,000 years ago is characterized by a series of founder events as modern humans expanded into multiple continents. Population genetics theory predicts an increase of mutational load in populations undergoing serial founder effects during range expansions. To test this hypothesis, we have sequenced full genomes and high-coverage exomes from 7 geographically divergent human populations from Namibia, Congo, Algeria, Pakistan, Cambodia, Siberia and Mexico. We find that individual genomes vary modestly in the overall number of predicted deleterious alleles. We show via spatially explicit simulations that the observed distribution of deleterious allele frequencies is consistent with the OOA dispersal, particularly under a model where deleterious mutations are recessive. We conclude that there is a strong signal of purifying selection at conserved genomic positions within Africa, but that many predicted deleterious mutations have evolved as if they were neutral during the expansion out of Africa. Under a model where selection is inversely related to dominance, we show that OOA populations are likely to have a higher mutation load due to increased allele frequencies of nearly neutral variants that are recessive or partially recessive.

Genomics

Approaches to estimating inbreeding coefficients in clinical isolates of Plasmodium falciparum from genomic sequence data

The advent of whole-genome sequencing has generated increased interest in modeling the structure of strain mixture within clinicial infections of Plasmodium falciparum (Pf). The life cycle of the parasite implies that the mixture of multiple strains within an infected individual is related to the out-crossing rate across populations, making methods for measuring this process in situ central to understanding the genetic epidemiology of the disease. In this paper, we show how to estimate inbreeding coefficients using genomic data from Pf clinical samples, providing a simple metric for assessing within-sample mixture that connects to an extensive literature in population genetics and conservation ecology. Features of the P. falciparum genome mean that some standard methods for inbreeding coefficients and related F-statistics cannot be used directly. Here, we review an initial effort to estimate the inbreeding coefficient within clinical isolates of P. falciparum and provide several generalizations using both frequentist and Bayesian approaches. The Bayesian approach connects these estimates to the Balding-Nichols model, a mainstay within genetic epidemiology. We provide simulation results on the performance of the estimators and show their use on ~ 1500 samples from the PF3K data set. We also compare the results to output from a recent mixture model for within-sample strain mixture, showing that inbreeding coefficients provide a strong proxy for the results of these more complex models. We provide the methods described within an open-source R package pfmix.

Genomics

Genome variation and meiotic recombination in Plasmodium falciparum: insights from deep sequencing of genetic crosses

The malaria parasite Plasmodium falciparum has a great capacity for evolutionary adaptation to evade host immunity and develop drug resistance. Current understanding of parasite evolution is impeded by the fact that a large fraction of the genome is either highly repetitive or highly variable, and thus difficult to analyse using short read technologies. Here we describe a resource of deep sequencing data on parents and progeny from genetic crosses, which has enabled us to perform the first genome-wide, integrated analysis of SNP, INDEL and complex polymorphisms, using Mendelian error rates as an indicator of genotypic accuracy. These data reveal that INDELs are exceptionally abundant, being more common than SNPs and thus the dominant mode of polymorphism within the core genome. We use the high density of SNP and INDEL markers to analyse patterns of meiotic recombination, confirming a high rate of crossover events, and providing the first estimates for the rate of non-crossover events and the length of conversion tracts. We observe several instances of recombination that modify copy number variants associated with drug resistance, demonstrating a mechanism whereby fitness costs associated with resistance mutations could be compensated and greater phenotypic plasticity could be acquired. We describe a novel web application that allows these data to be explored in detail.

Genomics

A genomic region containing RNF212 and CPLX1 is associated with sexually-dimorphic recombination rate variation in wild Soay sheep (Ovis aries).

Meiotic recombination breaks down linkage disequilibrium and forms new haplotypes, meaning thatit is an important driver of diversity in eukaryotic genomes. Understanding the causes of variation in recombination rate is important in interpreting and predicting evolutionary phenomena and forunderstanding the potential of a population to respond to selection. However, despite attention inmodel systems, there remains little data on how recombination rate varies at the individual level in natural populations. Here, we used extensive pedigree and high-density SNP information in a wild population of Soay sheep (Ovis aries) to investigate the genetic architecture of individual autosomal recombination rate. Individual rates were high relative to other mammal systems, and were higher in males than in females (autosomal map lengths of 3748 cM and 2860 cM, respectively). The heritability of autosomal recombination rate was low but significant in both sexes(h2 = 0.16 & 0.12 in females and males, respectively). In females, 46.7% of the heritable variation was explained by a sub-telomeric region on chromosome 6; a genome-wide association study showed the strongest associations at the locus RNF212, with further associations observed at a nearby ~374kb region of complete linkage disequilibrium containing three additional candidate loci, CPLX1, GAK and PCGF3. A second region on chromosome 7 containing REC8 and RNF212B explained 26.2% of the heritable variation in recombination rate in both sexes. Comparative analyses with 40 other sheep breeds showed that haplotypes associated with recombination rates are both old and globally distributed. Both regions have been implicated in rate variation in mice, cattle and humans, suggesting a common genetic architecture of recombination rate variation in mammals.\n\nAUTHOR SUMMARYRecombination offers an escape from genetic linkage by forming new combinations of alleles, increasing the potential for populations to respond to selection. Understanding the causes and consequences of individual recombination rates are important in studies of evolution and genetic improvement, yet little is known on how rates vary in natural systems. Using data from a wild population of Soay sheep, we show that individual recombination rate is heritable and differs between the sexes, with the majority of genetic variation in females explained by a genomic region containing thegenes RNF212 and CPLX1.

Genomics

Whole genome sequencing of field isolates reveals extensive genetic diversity in Plasmodium vivax from Colombia

Plasmodium vivax is the most prevalent malarial species in South America and exerts a substantial burden on the populations it affects. The control and eventual elimination of P. vivax are global health priorities. Genomic research contributes to this objective by improving our understanding of the biology of P. vivax and through the development of new genetic markers that can be used to monitor efforts to reduce malaria transmission.\n\nHere we analyze whole-genome data from eight field samples from a region in Cordoba, Colombia where malaria is endemic. We find considerable genetic diversity within this population, a result that contrasts with earlier studies suggesting that P. vivax had limited diversity in the Americas. We also identify a selective sweep around a substitution known to confer resistance to sulphadoxine-pyrimethamine (SP). This is the first observation of a selective sweep for SP resistance in this species. These results indicate that P. vivax has been exposed to SP pressure even when the drug is not in use as a first line treatment for patients afflicted by this parasite. We identify multiple non-synonymous substitutions in three other genes known to be involved with drug resistance in Plasmodium species. Finally, we found extensive microsatellite polymorphisms. Using this information we developed 18 polymorphic and easy to score microsatellite loci that can be used in epidemiological investigations in South America.\n\nAuthor SummaryAlthough P. vivax is not as deadly as the more widely studied P. falciparum, it remains a pressing global health problem. Here we report the results of a whole-genome study of P. vivax from Cordoba, Colombia, in South America. This parasite is the most prevalent in this region. We show that the parasite population is genetically diverse, which is contrary to expectations from earlier studies from the Americas. We also find molecular evidence that resistance to an anti-malarial drug has arisen recently in this region. This selective sweep indicates that the parasite has been exposed to a drug that is not used as first-line treatment for this malaria parasite. In addition to extensive single nucleotide and microsatellite polymorphism, we report 18 new genetic loci that might be helpful for fine-scale studies of this species in the Americas.

Genomics

Determinants of RNA metabolism in the Schizosaccharomyces pombe genome

To decrypt the regulatory code of the genome, sequence elements must be defined that determine the kinetics of RNA metabolism and thus gene expression. Here we attempt such decryption in an eukaryotic model organism, the fission yeast S. pombe. We first derive an improved genome annotation that redefines borders of 36% of expressed mRNAs and adds 487 non-coding RNAs (ncRNAs). We then combine RNA labeling in vivo with mathematical modeling to obtain rates of RNA synthesis and degradation for 5,484 expressed RNAs and splicing rates for 4,958 introns. We identify functional sequence elements in DNA and RNA that control RNA metabolic rates, and quantify the contributions of individual nucleotides to RNA synthesis, splicing, and degradation. Our approach reveals distinct kinetics of mRNA and ncRNA metabolism, separates antisense regulation by transcription interference from RNA interference, and provides a general tool for studying the regulatory code of genomes.

Genomics

Tying down loose ends in the Chlamydomonas genome

The Chlamydomonas genome has been sequenced, assembled and annotated to produce a rich resource for genetics and molecular biology in this well-studied model organism. The annotated genome is very rich in open reading frames upstream of the annotated coding sequence ( uORFs): almost three quarters of the assigned transcripts have at least one uORF, and frequently more than one. This is problematic with respect to the standard scanning model for eukaryotic translation initiation. These uORFs can be grouped into three classes: class 1, initiating in-frame with the coding sequence (cds) (thus providing a potential in-frame N-terminal extension); class 2, initiating in the 5UT and terminating out-of-frame in the cds; and class 3, initiating and terminating within the 5UT. Multiple bioinformatics criteria (including analysis of Kozak consensus sequence agreement and BLASTP comparisons to the closely related Volvox genome, and statistical comparison to cds and to random-sequence controls) indicate that of ~4000 class 1 uORFs, approximately half are likely in vivo translation initiation sites. The proposed resulting N-terminal extensions in many cases will sharply alter the predicted biochemical properties of the encoded proteins. These results suggest significant modifications in ~2000 of the ~20,000 transcript models with respect to translation initiation and encoded peptides. In contrast, class 2 uORFs may be subject to purifying selection, and the existent ones (surviving selection) are likely inefficiently translated. Class 3 uORFs are remarkably similar to random sequence expectations with respect to size, number and composition and therefore may be largely selectively neutral; their very high abundance (found in more than half of transcripts, frequently with multiple uORFs per transcript) nevertheless suggests the possibility of translational regulation on a wide scale.

Genomics

Genome-wide association and prediction reveals the genetic architecture of cassava mosaic disease resistance and prospects for rapid genetic improvement

Cassava (Manihot esculenta) is a crucial, under-researched crop feeding millions worldwide, especially in Africa. Cassava mosaic disease (CMD) has plagued production in Africa for over a century. Bi-parental mapping studies suggest primarily a single major gene mediates resistance. To be certain and to potentially identify new loci we conducted the first genome-wide association mapping study in cassava with 6128 African breeding lines. We also assessed the accuracy of genomic selection to improve CMD resistance. We found a single region on chromosome 8 accounts for most resistance but also identified 13 small effect regions. We found evidence that two epistatic loci and/or alternatively multiple resistance alleles exist at major QTL. We identified two peroxidases and one thioredoxin as candidate genes. Genomic prediction of additive and total genetic merit was accurate for CMD and will be effective both for selecting parents and identifying highly resistant clones as varieties.

Genomics

Centralizing content and distributing labor: a community model for curating the very long tail of microbial genomes.

The last 20 years of advancement in DNA sequencing technologies have led to the sequencing of thousands of microbial genomes, creating mountains of genetic data. While our efficiency in generating the data improves almost daily, applying meaningful relationships between the taxonomic and genetic entities requires a new approach. Currently, the knowledge is distributed across a fragmented landscape of resources from government-funded institutions such as NCBI and Uniprot to topic-focused databases like the ODB3 database of prokaryotic operons, to the supplemental table of a primary publication. A major drawback to large scale, expert curated databases is the expense of maintaining and extending them over time. No entity apart from a major institution with stable long-term funding can consider this, and their scope is limited considering the magnitude of microbial data being generated daily. Wikidata is an, openly editable, semantic web compatible framework for knowledge representation. Its a project of the Wikimedia Foundation and offers knowledge integration capabilities ideally suited to the challenge of representing the exploding body of information about microbial genomics. We are developing a microbial specific data model, based on Wikidatas semantic web compatibility, that represents bacterial species, strains and the gene and gene products that define them. Currently, we have loaded 1736 gene items and 1741 protein items for two strains of the human pathogenic bacteria Chlamydia trachomatis and used this subset of data as an example of the empowering utility of this model. In our next phase of development we will expand by adding another 118 bacterial genomes and their gene and gene products, totaling over ~900,000 additional entities. This aggregation of knowledge will be a platform for community-driven collaboration, allowing the networking of microbial genetic data through the sharing of knowledge by both the data and domain expert.

Genomics

Shared genomic variants: identification of transmission routes using pathogen deep sequence data

Sequencing pathogen samples during a communicable disease outbreak is becoming an increasingly common procedure in epidemiological investigations. Identifying who infected whom sheds considerable light on transmission patterns, high-risk settings and subpopulations, and infection control effectiveness. Genomic data shed new light on transmission dynamics, and can be used to identify clusters of individuals likely to be linked by direct transmission. However, identification of individual routes of infection via single genome samples typically remains uncertain. Here, we investigate the potential of deep sequence data to provide greater resolution on transmission routes, via the identification of shared genomic variants. We assess several easily implemented methods to identify transmission routes using both shared variants and genetic distance, demonstrating that shared variants can provide considerable additional information in most scenarios. While shared variant approaches identify relatively few links in the presence of a small transmission bottleneck, these links are highly confident. Furthermore, we proposed hybrid approach additionally incorporating phylogenetic distance to provide greater resolution. We apply our methods to data collected during the 2014 Ebola outbreak, identifying several likely routes of transmission. Our study highlights the power of pathogen deep sequence data as a component of outbreak investigation and epidemiological analyses.

Genomics

De novo Genome Assembly of Geosmithia morbida, the Causal Agent of Thousand Cankers Disease

Background: Geosmithia morbida is a filamentous ascomycete that causes Thousand Cankers Disease in the eastern black walnut tree. This pathogen is commonly found in the western U.S.; however, recently the disease was also detected in several eastern states where the black walnut lumber industry is concentrated. G. morbida is one of two known phytopathogens within the genus Geosmithia, and it is vectored into the host tree via the walnut twig beetle.\n\nResults: We present the first de novo draft genome of G. morbida. It is 26.5 Mbp in length and contains less than 1% repetitive elements. The genome possesses an estimated 6,273 genes, 277 of which are predicted to encode proteins with unknown functions. Approximately 31.5% of the proteins in G. morbida are homologous to proteins involved in pathogenicity, and 5.6% of the proteins contain signal peptides that indicate these proteins are secreted.\n\nConclusions: Several studies have investigated the evolution of pathogenicity in pathogens of agricultural crops; forest fungal pathogens are often neglected because research efforts are focused on food crops. G. morbida is one of the few tree phytopathogens to be sequenced, assembled and annotated. The first draft genome of G. morbida serves as a valuable tool for comprehending the underlying molecular and evolutionary mechanisms behind pathogenesis within the Geosmithia genus.

Genomics

Genome-wide histone modification patterns in Kluyveromyces Lactis reveal evolutionary adaptation of a heterochromatin-associated mark

The packaging of eukaryotic genomes into nucleosomes plays critical roles in all DNA-templated processes, and chromatin structure has been implicated as a key factor in the evolution of gene regulatory programs. While the functions of many histone modifications appear to be highly conserved throughout evolution, some well-studied modifications such as H3K9 and H3K27 methylation are not found in major model organisms such as Saccharomyces cerevisiae, while other modifications gain/lose regulatory functions during evolution. To study such a transition we focused on H3K9 methylation, a heterochromatin mark found in metazoans and in the fission yeast S. pombe, but which has been lost in the lineage leading to the model budding yeast S. cerevisiae. We show that this mark is present in the relatively understudied yeast Kluyveromyces lactis, a Hemiascomycete that diverged from S. cerevisiae prior to the whole-genome duplication event that played a key role in the evolution of a primarily fermentative lifestyle. We mapped genome-wide patterns of H3K9 methylation as well as several conserved modifications. We find that well-studied modifications such as H3K4me3, H3K36me3, and H3S10ph exhibit generally conserved localization patterns. Interestingly, we show H3K9 methylation in K. lactis primarily occurs over highly-transcribed regions, including both Pol2 and Pol3 transcription units. We identified the H3K9 methylase as the ortholog of Set6, whose function in S. cerevisiae is obscure. Functionally, we show that deletion of KlSet6 does not affect highly H3K9me3-marked genes, providing another example of a major disconnect between histone mark localization and function. Together, these results shed light on surprising plasticity in the function of a widespread chromatin mark.

Genomics