Search bioRxivSearch

Biology subjects

Howison, M.

Publications and source records attributed to Howison, M..

3 recordsLinked to original sources

Measurement error and variant-calling in deep Illumina sequencing of HIV

MotivationNext-generation deep sequencing of viral genomes, particularly on the Illumina platform, is increasingly applied in HIV research. Yet, there is no standard protocol or method used by the research community to account for measurement errors that arise during sample preparation and sequencing. Correctly calling high and low frequency variants while controlling for erroneous variant calls is an important precursor to downstream interpretation, such as studying the emergence of HIV drug-resistance mutations, which in turn has clinical applications and can improve patient care.\n\nResultsWe developed a new variant-calling pipeline, hivmmer, for Illumina sequences from HIV viral genomes. First, we validated hivmmer by comparing it to other variant-calling pipelines on real HIV plasmid data sets, which have known sequences. We found that hivmmer achieves a lower rate of erroneous variant calls, and that all methods agree on the frequency of correctly called variants. Next, we compared the methods on an HIV plasmid data set that was sequenced using an amplicon-tagging protocol called Primer ID, which is designed to reduce errors and amplification bias during library preparation. We show that the Primer ID consensus does indeed have fewer erroneous variant calls compared to the variant-calling pipelines, and that hivmmer more closely approaches this low error rate compared to the other pipelines. Surprisingly, the frequency estimates from the Primer ID consensus do not differ significantly from those of the variant-calling pipelines. Finally, we built a predictive model for classifying errors in the hivmmer alignment, and show that it achieves high accuracy for identifying erroneous variant calls.\n\nAvailabilityhivmmer is freely available for non-commercial use from https://github.com/mhowison/hivmmer.\n\nContactmhowison@brown.edu

bioinformatics

Improved phylogenetic resolution within Siphonophora (Cnidaria) with implications for trait evolution

Siphonophores are a diverse group of hydrozoans (Cnidaria) that are found at all depths of the ocean - from the surface, like the familiar Portuguese man of war, to the deep sea. Siphonophores play an important role in ocean ecosystems, and are among the most abundant gelatinous predators. A previous phylogenetic study based on two ribosomal RNA genes provided insight into the internal relationships between major siphonophore groups, however there was little support for many deep relationships within the clade Codonophora. Here, we present a new siphonophore phylogeny based on new transcriptome data from 30 siphonophore species analyzed in combination with 13 publicly available genomic and transcriptomic datasets. We use this new phylogeny to reconstruct several traits that are central to siphonophore biology, including sexual system (monoecy vs. dioecy), gain and loss of zooid types, life history traits, and habitat. The phylogenetic relationships in this study are largely consistent with the previous phylogeny, but we find strong support for new clades within Codonophora that were previously unresolved. These results have important implications for trait evolution within Siphonophora, including favoring the hypothesis that monoecy arose twice.

evolutionary biology

Revising transcriptome assemblies with phylogenetic information in Agalma1.0

MotivationOne of the most common transcriptome assembly errors is to mistake different transcripts of the same gene as transcripts from multiple closely related genes. It is difficult to identify these errors during assembly, but in a phylogenetic analysis these errors can be diagnosed from gene trees containing clades of tips from the same species with improbably short branch lengths.\n\nResultstreeinform is a module implemented in Agalma1.0 that uses phylogenetic analyses across species to refine transcriptome assemblies. It identifies transcripts of the same gene that were incorrectly assigned to multiple genes and reassign them as transcripts of the same gene.\n\nAvailability and Implementationtreeinform is implemented in Agalma1.0, available at https://bitbucket.org/caseywdunn/agalma.\n\nContactaugust_guang@brown.edu\n\nSupplementary informationSupplementary information is available at bioRxiv.

bioinformatics