Search bioRxiv⌕ Search

Biology subjects

Fodor, A.

Publications and source records attributed to Fodor, A..

2 recordsLinked to original sources

Systematic classification error profoundly impacts inference in high-depth Whole Genome Shotgun Sequencing datasets

There is little consensus in the literature as to which approach for classification of Whole Genome Shotgun (WGS) sequences is best. In this paper, we examine two of the most popular algorithms, Kraken2 and Metaphlan2 utilizing four publicly available datasets. As expected from previous literature, we found that Kraken2 reports more overall taxa while Metaphlan2 reports fewer taxa while classifying fewer overall reads. To our surprise, however, Kraken 2 reported not only more taxa but many more taxa that were significantly associated with metadata. This implies that either Kraken2 is more sensitive to taxa that are biologically relevant and are simply missed by Metaphlan2, or that Kraken2s classification errors are generated in such a way to impact inference. To discriminate between these two possibilities, we compared Spearman correlations coefficients of each taxa against each taxa with higher abundance from the same dataset. We found that Kraken2, but not Metaphlan2, showed a consistent pattern of classifying low abundance taxa that generated high correlation coefficients with higher abundance taxa. Neither Metaphlan2, nor 16S sequences that were available for two of our four datasets, showed this pattern. Simple simulations based on a variable Poisson error rate sampled from the uniform distribution with an average error rate of 0.0005 showed strikingly strong concordance with the observed correlation patterns from Kraken2. Our results suggest that Kraken2 consistently misclassifies high abundance taxa into the same erroneous low abundance taxa creating "phantom" taxa have a similar pattern of inference as the high abundance source. Because of the large sequencing depths of modern WGS cohorts, these "phantom" taxa will appear statistically significant in statistical models even with a low overall rate of classification error from Kraken. Our simulations suggest that this can occur with average error rates as low as 1 in 2,000 reads. These data suggest a novel metric for evaluating classifier accuracy and suggest that the pattern of classification errors should be considered in addition to overall classification error rate since consistent classification errors have a more profound impact on inference compared to classification errors that do not always result in assignment to the same erroneous taxa. This work highlights fundamental questions on how classifiers function and interact with large sequencing depth and statistical models that still need to be resolved for WGS, especially if correlation coefficients between taxa are to be used to build covariance networks. Our work also suggests that despite its limitations, 16S rRNA sequencing may still be useful as neither of the two most popular 16S classifiers showed these patterns of inflated correlation coefficients between taxa.

bioinformatics↗

Genome-wide association and prediction studies using a grapevine diversity panel give insights into the genetic architecture of several traits of interest

To cope with the challenges faced by agriculture, speeding-up breeding programs is a worthy endeavor, especially for perennials such as grapevine, but requires understanding the genetic architecture of target traits. To go beyond the mapping of quantitative trait locus (QTL) in bi-parental crosses, we exploited a diverse panel of 279 Vitis vinifera L. cultivars. This panel planted in five blocks in the vineyard was phenotyped over several years for 127 traits including yield components, organic acids, aroma precursors, polyphenols, and a water stress indicator. The panel was genotyped for 63k single nucleotide polymorphisms (SNPs) by combining an 18K microarray and genotyping-by-sequencing (GBS). The experimental design allowed to reliably assess the genotypic values for most traits. Marker densification via GBS markedly increased the proportion of genetic variance explained by SNPs, and two multi-SNP models identified QTLs not found by a SNP-by-SNP model. Overall, 489 reliable QTLs were detected for 41% more response variables than by a SNP-by-SNP model with microarray-only SNPs, many new ones compared to the results from bi-parental crosses. Prediction accuracy higher than 0.42 was obtained for 50% of the response variables. Our overall approach as well as QTL and prediction results provide insights into the genetic architecture of target traits. New candidate genes and the application in breeding are discussed.

genetics↗