Search bioRxivSearch

Biology subjects

Brant C Faircloth

Publications and source records attributed to Brant C Faircloth.

5 recordsLinked to original sources

Identifying Conserved Genomic Elements and Designing Universal Probe Sets To Enrich Them

Targeted enrichment of conserved genomic regions is a popular method for collecting large amounts of sequence data from non-model taxa for phylogenetic, phylogeographic, and population genetic studies. Yet, few open-source workflows are available to identify conserved genomic elements shared among divergent taxa and to design enrichment baits targeting these regions. These shortcomings limit the application of targeted enrichment methods to many organismal groups. Here, I describe a universal workflow for identifying conserved genomic regions in available genomic data and for designing targeted enrichment baits to collect data from these conserved regions. I demonstrate how this computational approach can be applied to diverse organismal groups by identifying sets of conserved loci and designing enrichment baits targeting thousands of these loci in the understudied arthropod groups Arachnida, Coleoptera, Diptera, Hemiptera, or Lepidoptera. I then use in silico analyses to demonstrate that these conserved loci reconstruct the accepted relationships among genome sequences from the focal arthropod orders, and we perform in vitro validation of the Arachnid probe set as part of a separate manuscript (Starrett et al. Submitted). All of the documentation, design steps, software code, and probe sets developed here are available under an open-source license for restriction-free testing and use by any research group, and although the examples in this manuscript focus on understudied and exceptionally diverse arthropod groups, the software workflow is applicable to all organismal groups having some form of pre-existing genomic information.

Genomics

High Phylogenetic Utility of an Ultraconserved Element Probe Set Designed for Arachnida

Arachnida is an ancient, diverse, and ecologically important animal group that contains a number of species of interest for medical, agricultural, and engineering applications. Despite this applied importance, many aspects of the arachnid tree of life remain unresolved, hindering comparative approaches to arachnid biology. Biologists have made considerable efforts to resolve the arachnid phylogeny; yet, limited and challenging morphological characters, as well as a dearth of genetic resources, have confounded these attempts. Here, we present a genomic toolkit for arachnids featuring hundreds of conserved DNA regions (ultraconserved elements or UCEs) that allow targeted sequencing of any species in the arachnid tree of life. We used recently developed capture probes designed from conserved genomic regions of available arachnid genomes to enrich a sample of loci from 32 diverse arachnids. Sequence capture returned an average of 487 UCE loci for all species, with a range from 170 to 722. Phylogenetic analysis of these UCEs produced a highly resolved arachnid tree with relationships largely consistent with recent transcriptome-based phylogenies. We also tested the phylogenetic informativeness of UCE probes within the spider, scorpion, and harvestman orders, demonstrating the utility of these markers at shallower taxonomic scales, even down to the level of species differences. This probe set will open the door to phylogenomic and population genomic studies across the arachnid tree of life, enabling systematics, species delimitation, species discovery, and conservation of these diverse arthropods.

Genomics

Adapterama IV: Sequence Capture of Dual-digest RADseq Libraries with Identifiable Duplicates (RADcap)

Molecular ecologists seek to genotype hundreds to thousands of loci from hundreds to thousands of individuals at minimal cost per sample. Current methods such as restriction site associated DNA sequencing (RADseq) and sequence capture are constrained by costs associated with inefficient use of sequencing data and sample preparation, respectively. Here, we demonstrate RADcap, an approach that combines the major benefits of RADseq (low cost with specific start positions) with those of sequence capture (repeatable sequencing of specific loci) to significantly increase efficiency and reduce costs relative to current approaches. The RADcap approach uses a new version of dual-digest RADseq (3RAD) to identify candidate SNP loci for capture bait design, and subsequently uses custom sequence capture baits to consistently enrich candidate SNP loci across many individuals. We combined this approach with a new library preparation method for identifying and removing PCR duplicates from 3RAD libraries, which allows researchers to process RADseq data using traditional pipelines, and we tested the RADcap method by genotyping sets of 96 to 384 Wisteria plants. Our results demonstrate that our RADcap method: 1) can methodologically reduce (to <5%) and computationally remove PCR duplicate reads from data; (2) achieves 80-90% reads-on-target in 11 of 12 enrichments; (3) returns consistent coverage ([&ge;]4x) across >90% of individuals at up to 99.9% of the targeted loci; (4) produces consistently high occupancy matrices of genotypes across hundreds of individuals; and (5) is inexpensive, with reagent and sequencing costs totaling <$6/sample and adapter and primer costs of only a few hundred dollars.

Genetics

PHYLUCE is a software package for the analysis of conserved genomic loci

SummaryTargeted enrichment of conserved and ultraconserved genomic elements allows universal collection of phylogenomic data from hundreds of species at multiple time scales (< 5 Ma to > 300 Ma). Prior to downstream inference, data from these types of targeted enrichment studies must undergo pre-processing to assemble contigs from sequence data; identify targeted, enriched loci from the off-target background data; align enriched contigs representing conserved loci to one another; and prepare and manipulate these alignments for subsequent phylogenomic inference. PHYLUCE is an efficient and easy-to-install software package that accomplishes these tasks across hundreds of taxa and thousands of enriched loci.\n\nAvailability and ImplementationPHYLUCE is written for Python 2.7. PHYLUCE is supported on OSX and Linux (RedHat/CentOS) operating systems. PHYLUCE source code is distributed under a BSD-style license from https://www.github.com/fairclothUlab/phyluce/. PHYLUCE is also available as a package (https://binstar.org/fairclothUlab/phyluce) for the Anaconda Python distribution that installs all dependencies, and users can request a PHYLUCE instance on iPlant Atmosphere (tag: phyluce). The software manual and a tutorial are available from http://phyluce.readthedocs.org/en/latest/ and test data are available from doi: 10.6084/m9.figshare.1284521.\n\nContactbrant@fairclothUlab.org\n\nSupplementary informationSupplementary Figure 1.\n\nO_FIG O_LINKSMALLFIG WIDTH=178 HEIGHT=200 SRC=\"FIGDIR/small/027904_figS1.gif\" ALT=\"Figure 1\">\nView larger version (40K):\norg.highwire.dtl.DTLVardef@1f6fb2aorg.highwire.dtl.DTLVardef@1e39afforg.highwire.dtl.DTLVardef@1d51cf1org.highwire.dtl.DTLVardef@5f2eda_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOSupplementary Figure 1.C_FLOATNO PHYLUCE workflow for phylogenomic analyses of data collected from conserved genomic loci using targeted enrichment.\n\nC_FIG

Bioinformatics

Sequence capture of ultraconserved elements from bird museum specimens

New DNA sequencing technologies are allowing researchers to explore the genomes of the millions of natural history specimens collected prior to the molecular era. Yet, we know little about how well specific next-generation sequencing (NGS) techniques work with the degraded DNA typically extracted from museum specimens. Here, we use one type of NGS approach, sequence capture of ultraconserved elements (UCEs), to collect data from bird museum specimens as old as 120 years. We targeted approximately 5,000 UCE loci in 27 Western Scrub-Jays (Aphelocoma californica) representing three evolutionary lineages, and we collected an average of 3,749 UCE loci containing 4,460 single nucleotide polymorphisms (SNPs). Despite older specimens producing fewer and shorter loci in general, we collected thousands of markers from even the oldest specimens. More sequencing reads per individual helped to boost the number of UCE loci we recovered from older specimens, but more sequencing was not as successful at increasing the length of loci. We detected contamination in some samples and determined contamination was more prevalent in older samples that were subject to less sequencing. For the phylogeny generated from concatenated UCE loci, contamination led to incorrect placement of some individuals. In contrast, a species tree constructed from SNPs called within UCE loci correctly placed individuals into three monophyletic groups, perhaps because of the stricter analytical procedures we used for SNP calling. This study and other recent studies on the genomics of museums specimens have profound implications for natural history collections, where millions of older specimens should now be considered genomic resources.

Evolutionary Biology