Search bioRxivSearch

Biology subjects

Creevey, C. J.

Publications and source records attributed to Creevey, C. J..

3 recordsLinked to original sources

Computational haplotype recovery and long-read validation identifies novel isoforms of industrially relevant enzymes from natural microbial communities

Elucidation of population-level diversity of microbiomes is a significant step towards a complete understanding of the evolutionary, ecological and functional importance of microbial communities. Characterizing this diversity requires the recovery of the exact DNA sequence (haplotype) of each gene isoform from every individual present in the community. To address this, we present Hansel and Gretel: a freely-available data structure and algorithm, providing a software package that reconstructs the most likely haplotypes from metagenomes. We demonstrate recovery of haplotypes from short-read Illumina data for a bovine rumen microbiome, and verify our predictions are 100% accurate with long-read PacBio CCS sequencing. We show that Gretels haplotypes can be analyzed to determine a significant difference in mutation rates between core and accessory gene families in an ovine rumen microbiome. All tools, documentation and data for evaluation are open source and available via our repository: https://github.com/samstudio8/gretel

bioinformatics

Probabilistic Recovery Of Cryptic Haplotypes From Metagenomic Data

The cryptic diversity of microbial communities represent an untapped biotechnological resource for biomining, biorefining and synthetic biology. Revealing this information requires the recovery of the exact sequence of DNA bases (or \"haplotype\") that constitutes the genes and genomes of every individual present. This is a computationally difficult problem complicated by the requirement for environmental sequencing approaches (metagenomics) due to the resistance of the constituent organisms to culturing in vitro.\n\nHaplotypes are identified by their unique combination of DNA variants. However, standard approaches for working with metagenomic data require simplifications that violate assumptions in the process of identifying such variation. Furthermore, current haplotyping methods lack objective mechanisms for choosing between alternative haplotype reconstructions from microbial communities.\n\nTo address this, we have developed a novel probabilistic approach for reconstructing haplotypes from complex microbial communities and propose the \"metahaplome\" as a definition for the set of haplotypes for any particular genomic region of interest within a metagenomic dataset. Implemented in the twin software tools Hansel and Gretel, the algorithm performs incremental probabilistic haplotype recovery using Naive Bayes -- an efficient and effective technique.\n\nOur approach is capable of reconstructing the haplotypes with the highest likelihoods from metagenomic datasets without a priori knowledge or making assumptions of the distribution or number of variants. Additionally, the algorithm is robust to sequencing and alignment error without altering or discarding observed variation and uses all available evidence from aligned reads. We validate our approach using synthetic metahaplomes constructed from sets of real genes, and demonstrate its capability using metagenomic data from a complex HIV-1 strain mix. The results show that the likelihood framework can allow recovery from microbial communities of cryptic functional isoforms of genes with 100% accuracy.

bioinformatics

Comparison and Characterisation of Mutation Calling from Whole Exome and RNA Sequencing Data for Liver and Muscle Tissue in Lactating Holstein Cows Divergent for Fertility

Whole exome sequencing has had low uptake in livestock species, despite allowing accurate analysis of single nucleotide variant (SNV) mutations. Transcriptomic data in the form of RNA sequencing has been generated for many livestock species and also represents a source of mutational information. However, there is little information on the accuracy of using this data for the identification of SNVs. We generated a bovine exome capture design and used it to sequence and call mutations from a lactating dairy cow model genetically divergent for fertility (Fert+, n=8; Fert-, n=8). We compared mutations called from liver and muscle transcriptomes from the same animals. Our exome capture demonstrated 99.1% coverage of the exome design of 56.7MB, whereas transcriptomes covered 55 and 46.5% of the exome, or 24.4 and 20.7MB, in liver and muscle respectively after filtering. We found that specificity of SNVs in the transcriptome data is approximately 75% following basic hard-filtering, and could be increased to above 80% by increasing the minimum threshold of reads covering SNVs, but this effect was negated in more highly covered SNVs. RNA-DNA differences, SNVs found in transcriptome but not exome, were discovered and shown to have significantly increased levels of transition mutations in both tissues. Functional annotation of non-synonymous SNVs specific to the high and low fertility phenotypes identified immune response-related genes, supporting previous work that has identified differential expression in the same genes. Publically available RNAseq data may be analysed in a similar way to further increase the utility of this resource.\n\nSummaryThe exome and transcriptome both relate to the same protein-coding regions of the genome. There has been sparse research on characterising mutations in RNA and DNA within the same individuals. Here we characterise the similarities in our Holstein dairy cow animal model. We offer practical and biological results indicating that RNA sequencing is a useful proxy of exome sequencing, itself shown to be applicable to this livestock species using a previously untested commercial application. This potentially unlocks public RNA sequencing data for further analysis, also indicating that RNA-DNA differences may associate with transcriptomic divergence.

genomics