Search bioRxiv⌕ Search

Biology subjects

Bard, N. W.

Publications and source records attributed to Bard, N. W..

2 recordsLinked to original sources

The Barcode Inference Pipeline (BIP): From Sequencer Output to DNA Barcodes

DNA barcoding involves the recovery of a DNA sequence for a target gene region from its source specimen. This process gains complexity when multiple sequences are recovered from a specimen, as is often the case when data are generated by high-throughput sequencers. This diversity can reflect both methodological artifacts (e.g., chimeras, PCR errors, sequencing errors, tag jumps) and real template diversity in the DNA extract (e.g., contamination, endosymbionts, NUMTs, parasites). To support analysis of the sequence data from three million specimens annually, the Centre for Biodiversity Genomics (CBG) has developed BIP, the Barcode Inference Pipeline. Compatible with all sequencing platforms, BIP processes .fastq files and returns both target DNA barcodes and non-target sequences. To generate results, BIP implements quality and size filtration, demultiplexing, primer trimming, chimera scanning, sequence error correction, OTU delineation, and sequence identification. When analysis targets the cytochrome c oxidase 1 (COI) barcode region, BIP also assigns each OTU to a known BIN or identifies its nearest neighbour BIN. As final output, BIP returns summary files ready for upload to BOLD or for other downstream analyses. They include a taxonomic assignment for each OTU, generated by comparison with a DNA barcode reference library. We describe BIPs flexibility and structure, then demonstrate its functionality by analyzing COI sequence data from 100K specimens. Because of its capacity to disentangle target and non-target sequences, BIP outperforms an alternative software package, ONTbarcoder, in several important ways. To ease access, installation, and functionality, BIP is provided as a Docker container (github.com/cbg-innov/BIP).

bioinformatics↗

The Metabarcoding Analysis Pipeline (MAP): Simple, accurate, and flexible metabarcoding

Current metabarcoding pipelines are inflexible with respect to study design and are poorly suited to long-read sequence data. To address these limitations, we developed MAP, the Metabarcoding Analysis Pipeline, which is a sequence-to-answer workflow supporting the analysis of amplicons from highly multiplexed and replicated study designs. Although MAP can analyze amplicons of any length from any genetic marker, it includes several features tailored to long-read COI metabarcoding. MAP installs from a Docker container and requires only sequence data, a parameters file, and a reference library. It produces intuitive reports, enabling users to evaluate their data immediately after analysis. We validate MAP by showing that it generates biodiversity estimates that correspond closely to a ground-truth dataset of single-specimen DNA barcode data and by demonstrating that it outperforms alternative platforms for COI metabarcoding. MAP is free, open-source, and available from: https://github.com/cbg-innov/MAP.

bioinformatics↗