Search bioRxivSearch

Biology subjects

Nasko, D. J.

Publications and source records attributed to Nasko, D. J..

2 recordsLinked to original sources

Fast and sensitive protein sequence homology searches using hierarchical cluster BLAST

The throughput of DNA sequencing continues to increase, allowing researchers to analyze genomes of interest at greater depths. An unintended consequence of this data deluge is the increased cost of analyzing these datasets. As a result, genome and metagenome annotation pipelines are left with a few options: (i) search against smaller reference databases, (ii) use faster, but less sensitive, algorithms to assess sequence similarities, or (iii) invest in computing hardware specifically designed to improve BLAST searches such as GPGPU systems and/or large CPU-rich clusters.\n\nWe present a pipeline that improves the speed of amino acid sequence homology searches with a minimal decrease in sensitivity and specificity by searching against hierarchical clusters. Briefly, the pipeline requires two homology searches: the first search is against a clustered version of the database and the second is against sequences belonging to clusters with a hit from the first search. We tested this method using two assembled viral metagenomes and three databases (Swiss-Prot, Metagenomes Online, and UniRef100). Hierarchical cluster homology searching proved to be 12-times faster than BLASTp and produced alignments that were nearly identical to BLASTp (precision=0.99; recall=0.97). This approach is ideal when searching large collections of sequences against large databases.

bioinformatics

RefSeq database growth influences the accuracy of k-mer-based species identification

Accurate species-level taxonomic classification and profiling of complex microbial communities remains a challenge due to homologous regions shared among closely related species and a sparse representation of non-human associated microbes in the database. Although the database undoubtedly has a strong influence on the sensitivity of taxonomic classifiers and profilers, to date, no study has carefully explored this topic on historical RefSeq releases and explored its impact on accuracy. In this study, we examined the influence of the database, over time, on k-mer based sequence classification and profiling. We present three major findings: (i) database growth over time resulted in more classified reads, but fewer species-level classifications and more species-level misclassifications; (ii) Bayesian re-estimation of abundance helped to recover species-level classifications when the exact target strain was present; and (iii) Bayesian reestimation struggled when the database lacked the target strain, resulting in a notable decrease in accuracy. In summary, our findings suggest that the growth of RefSeq over time has strongly influenced the accuracy of k-mer based classification and profiling methods, resulting in different classification results depending on the particular database used. These results suggest a need for new algorithms specially adapted for large genome collections and better measures of classification uncertainty.

bioinformatics