Search bioRxiv⌕ Search

Biology subjects

Hintikka, S.

Publications and source records attributed to Hintikka, S..

2 recordsLinked to original sources

The bacterial hitchhiker's guide to COI: Universal primer-based COI capture probes fail to exclude bacterial DNA, but 16S capture leaves metazoa behind

Environmental DNA (eDNA) metabarcoding from water samples has, in recent years, shown great promise for biodiversity monitoring. However, universal primers targeting the cytochrome oxidase I (COI) marker gene popular in metazoan studies have displayed high levels of nontarget amplification. To date, enrichment methods bypassing amplification have not been able to match the detection levels of conventional metabarcoding. This study evaluated the use of universal metabarcoding primers as capture probes to either isolate target DNA or to remove nontarget DNA, prior to amplification, by using biotinylated versions of universal metazoan and bacterial barcoding primers, namely metazoan COI (mlCOIintF) and bacterial 16S (515F). Additionally, each step of the protocol was assessed by amplifying for both metazoan COI (mlCOIintF/jgHCO2198) and bacterial 16S (515F/806R) to investigate the effect on the metazoan and bacterial communities. Bacterial read abundance increased significantly in response to the captures (COI library), while the quality of the captured DNA was also improved. The metazoan-based probe captured bacterial DNA in a range that was also amplifiable with the 16S primers, demonstrating the ability of universal capture probes to isolate larger fragments of DNA from eDNA. Although the use of the tested COI probe cannot be recommended for metazoan enrichment, based on the experimental results, the concept of capturing longer fragments could be applied to metazoan metabarcoding. By using a truly conserved site without a high-level taxonomic resolution as a target for capture, it may be possible to isolate DNA fragments large enough to span over a nearby barcoding region (e.g., COI), which can then be processed through a conventional metabarcoding-by-amplification protocol.

molecular biology↗

Bacteria are everywhere, even in your COI marker gene data!

The mitochondrial cytochrome C oxidase subunit I gene (COI) is commonly used in eDNA metabarcoding studies, especially for assessing metazoan diversity. Yet, a great number of COI operational taxonomic units or/and amplicon sequence variants are retrieved from such studies and referred to as "dark matter", and do not get a taxonomic assignment with a reference sequence. For a thorough investigation of this dark matter, we have developed the Dark mAtteR iNvestigator (DARN) software tool. A reference COI-oriented phylogenetic tree was built from 1,240 consensus sequences covering all the three domains of life, with more than 80% of those representing eukaryotic taxa. With respect to eukaryotes, consensus sequences at the family level were constructed from 183,330 retrieved from the Midori reference 2 database. Similarly, sequences from 559 bacterial genera and 41 archaeal were retrieved from the BOLD database. DARN makes use of the phylogenetic tree to investigate and quantify pre-processed sequences of amplicon samples to provide both a tabular and a graphical overview of phylogenetic assignments. To evaluate DARN, both environmental and bulk metabarcoding samples from different aquatic environments using various primer sets were analysed. We demonstrate that a large proportion of non-target prokaryotic organisms such as bacteria and archaea are also amplified in eDNA samples and we suggest bacterial COI sequences to be included in the reference databases used for the taxonomy assignment to allow for further analyses of dark matter. DARN source code is available on GitHub at https://github.com/hariszaf/darn and you may find it as a Docker at https://hub.docker.com/r/hariszaf/darn. Author summaryDARN is a software approach aiming to provide further insight in the COI amplicon data coming from environmental samples. Building a COI-oriented reference phylogeny tree is a challenging task especially considering the small number of microbial curated COI sequences deposited in reference databases; e.g ~4,000 bacterial and ~150 archaeal in BOLD. Apparently, as more and more such sequences are collated, the DARN approach improves. To provide a more interactive way of communicating both our approach and our results, we strongly suggest the reader to visit this Google Collab notebook where all steps are described step by step and also this GitHub page where our results are demonstrated. Our approach corroborates the known presence of microbial sequences in COI environmental sequencing samples and highlights the need for curated bacterial and archaeal COI sequences and their integration into reference databases (i.e. Midori, BOLD, etc). We argue that DARN will benefit researchers as a quality control tool for their sequenced samples in terms of distinguishing eukaryotic from non-eukaryotic OTUs/ASVs, but also in terms of understanding the unknown unknowns.

bioinformatics↗