Search bioRxivSearch

Biology subjects

Edgar, R. C.

Publications and source records attributed to Edgar, R. C..

9 recordsLinked to original sources

Alpha diversity metrics for noisy OTUs

Next-generation sequencing (NGS) of marker genes such as 16S ribosomal RNA is widely used to survey microbial communities. The in-sample (alpha) diversity of Operational Taxonomic Units (OTUs) is often summarized by metrics such as richness or entropy which are calculated from observed abundances, or by estimators such as Chao1 which extrapolate to unobserved OTUs. Most such measures are adopted from traditional biodiversity studies, where observational error can often be neglected. However, errors introduced by next-generation amplicon sequencing tend to induce spurious OTUs and spurious counts in OTU tables, both of which are especially prevalent at low abundances. In consequence, traditional metrics may be grossly inaccurate if they are naively applied to NGS OTU tables. In this work, we describe two novel alpha diversity estimators which are calculated from OTU abundances above a specified threshold. The singleton-free estimator (SFE) is a non-parametric estimator which is derived from a similar approach to Chao1 but extrapolates using doublet and triplet abundances rather than singletons and doublets. The octave estimator (OE) fits a log-normal distribution to non-singleton bars of an octave plot. We show that these estimators are effective under suitable conditions, but these conditions rarely apply in practice. We conclude that extrapolating to unobserved OTUs remains an open problem which is unlikely to be solved in the near future.

bioinformatics

Octave plots for visualizing diversity of microbial OTUs

Next-generation sequencing of marker genes such as 16S ribosomal RNA is widely used to survey microbial communities. The abundance distribution (AD) of Operational Taxonomic Units (OTUs) in a sample is typically summarized by alpha diversity metrics, e.g. richness and entropy, discarding information about the AD shape. In this work, we describe octave plots, histograms which visualize the shape of microbial ADs by binning on a logarithmic scale with base 2. Optionally, histogram bars are colored to indicate possible spurious OTUs due to sequence error and cross-talk. Octave plots enable assessment of (a) the shape and completeness of the distribution, (b) the effects of noise on measured diversity, (c) whether low-abundance OTUs should be discarded, (d) whether alpha diversity metrics and estimators are reliable, and (e) the additional sampling effort (i.e., read depth) required to obtain a complete census of the community. The utility of octave plots is illustrated in a re-analysis of a prostate cancer study showing that the reported core microbiome is most likely an artifact of experimental error.

bioinformatics

Taxonomy annotation errors in 16S rRNA and fungal ITS sequence databases

Sequencing of the 16S ribosomal RNA (rRNA) gene and the fungal Internal Transcribed Spacer (ITS) region is widely used to survey microbial communities. Specialized ribosomal sequence databases have been developed to support this approach including Greengenes, SILVA and RDP. Most taxonomy annotations in these databases are predictions from sequence rather than authoritative assignments based on studies of type strains or isolates. Here, I investigate the error rates of taxonomy annotations in these databases. I found 253,485 sequences with conflicting annotations in SILVA v128 and Greengenes v13.5 at ranks up to phylum (9,644 conflicts), indicating that the annotation error rate in these databases is ~15%. I found that 34% of non-singleton genera have overlapping subtrees in the Greengenes tree from 2001 according to the RDP taxonomy, most of which are probably due to branching order errors in the Greengenes tree, which is therefore an unreliable guide to phylogeny. Using a blinded test, I estimated that the annotation error rate of the RDP database is ~10%.

microbiology

Updating the 97% identity threshold for 16S ribosomal RNA OTUs

The 16S ribosomal RNA (rRNA) gene is widely used to survey microbial communities. Sequences are often clustered into Operational Taxonomic Units (OTUs) as proxies for species. The canonical clustering threshold is 97% identity, which was proposed in 1994 when few 16S rRNA sequences were available, motivating a reassessment on current data. Using a large set of high-quality 16S rRNA sequences from finished genomes, I assessed the correspondence of OTUs to species for five representative clustering algorithms using four accuracy metrics. All algorithms had comparable accuracy when tuned to a given metric. Optimal identity thresholds that best approximated species were [~]99% for full-length sequences and [~]100% for the V4 hypervariable region.

bioinformatics

UNBIAS: An attempt to correct abundance bias in 16S sequencing, with limited success

Next-generation amplicon sequencing of 16S ribosomal RNA is widely used to survey microbial communities. Alpha and beta diversities of these communities are often quantified on the basis of OTU frequencies in the reads. Read abundances are biased by factors including 16S copy number and PCR primer mismatches which can cause the read abundance distribution to diverge substantially from the species abundance distribution. Using mock community tests with species abundances determined independently by shotgun sequencing, I find that 16S amplicon read frequencies have no meaningful correlation with species frequencies (Pearson coefficient r close to zero). In addition, I show that that the Jaccard distance between the abundance distributions for reads of replicate samples, which ideally would be zero, is typically ~0.15 with values up to 0.71 for replicates sequenced in different runs. Using simulated communities, I estimate that the average rank of a dominant species in the reads is 3. I describe UNBIAS, a method that attempts to correct for abundance bias due to gene copy number and primer mismatches. I show that UNBIAS can achieve informative, but still poor, correlations (r ~0.6) between estimated and true abundances in the idealized case of mock samples where species are well known. However, r falls to ~0.4 when the closest reference species have 97% identity and to ~0.2 at 95% identity. This degradation is mostly explained by the increased difficulty in predicting 16S copy number when OTUs have lower similarity with the reference database, as will typically be the case in practice. 16S abundance bias therefore remains an unsolved problem, calling into question the naive use of alpha and beta diversity metrics based on frequency distributions.

bioinformatics

SINAPS: Prediction of microbial traits from marker gene sequences

Microbial communities are often studied by sequencing marker genes such as 16S ribosomal RNA. Marker gene sequences can be used to assess diversity and taxonomy, but do not directly measure functions arising from other genes in the community metagenome. Such functions can be predicted by algorithms that associate marker genes with experimentally determined traits in well-studied species. Typically, such methods use ancestral state reconstruction. Here I describe SINAPS, a new algorithm that predicts traits for marker gene sequences using a fast, simple word-counting algorithm that does not require alignments or trees. A measure of prediction confidence is obtained by bootstrapping. I tested SINAPS predictions from 16S V4 query sequences for traits including energy metabolism, Gram-positive staining, presence of a flagellum, V4 primer mismatches, and 16S copy number. Accuracy was >90% except for copy number, where a large majority of predictions were within +/-2 of the true value.

bioinformatics

UNCROSS: Filtering of high-frequency cross-talk in 16S amplicon reads

Next-generation amplicon sequencing is widely used for surveying biological diversity in applications such as microbial metagenomics, immune system repertoire analysis and targeted tumor sequencing of cancer-associated genes. In such studies, assignment of reads to incorrect samples (cross-talk) is a well-documented problem that is rarely considered in practice. By considering unexpected OTUs in artificial (mock) samples, I estimate that cross-talk occurred for ~2% of the reads in one Illumina GAIIx run and eleven Illumina MiSeq runs targeting 16S ribosomal RNA. I also describe UNCROSS, an algorithm for detecting and filtering cross-talk in OTU tables.

bioinformatics

UNOISE2: improved error-correction for Illumina 16S and ITS amplicon sequencing

Amplicon sequencing of tags such as 16S and ITS ribosomal RNA is a popular method for investigating microbial populations. In such experiments, sequence errors caused by PCR and sequencing are difficult to distinguish from true biological variation. I describe UNOISE2, an updated version of the UNOISE algorithm for denoising (error-correcting) Illumina amplicon reads and show that it has comparable or better accuracy than DADA2.

bioinformatics