Search bioRxivSearch

Biology subjects

Joanna L Kelley

Publications and source records attributed to Joanna L Kelley.

4 recordsLinked to original sources

Inferring local movement of pathogen vectors among hosts

Herbivores often move among spatially interspersed host plants, tracking high-quality resources through space and time. This dispersal is of particular interest for vectors of plant pathogens. Existing molecular tools to track such movement have yielded important insights, but often provide insufficient genetic resolution to infer spread at finer spatiotemporal scales. Here, we explore the use of Nextera-tagmented reductively-amplified DNA (NextRAD) sequencing to infer movement of a highly-mobile winged insect, the potato psyllid (Bactericera cockerelli), among host plants. The psyllid vectors the pathogen that causes zebra chip disease in potato (Solanum tuberosum), but understanding and managing the spread of this pathogen is limited by uncertainty about the insects host plant(s) outside of the growing season. We identified 8,443 polymorphic loci among psyllids separated spatiotemporally on potato or in patches of bittersweet nightshade (S. dulcumara), a weedy plant proposed to be the source of potato-colonizing psyllids. A subset of the psyllids on potato exhibited close genetic similarity to insects on nightshade, consistent with regular movement between these two host plants. However, a second subset of potato-collected psyllids was genetically distinct from those collected on bittersweet nightshade; this suggests that a currently unrecognized host-plant species could be contributing to psyllid populations in potato. Oftentimes, dispersal of vectors of plant or animal pathogens must be tracked at a relatively fine scale in order to understand, predict, and manage disease spread. We demonstrate that emerging sequencing technologies that detect SNPs across a vectors entire genome can be used to infer such localized movement.

Ecology

GBStools: A Unified Approach for Reduced Representation Sequencing and Genotyping

Reduced representation sequencing methods such as genotyping-by-sequencing (GBS) enable low-cost measurement of genetic variation without the need for a reference genome assembly. These methods are widely used in genetic mapping and population genetics studies, especially with non-model organisms. Variant calling error rates, however, are higher in GBS than in standard sequencing, in particular due to restriction site polymorphisms, and few computational tools exist that specifically model and correct these errors. We developed a statistical method to remove errors caused by restriction site polymorphisms, implemented in the software package GBStools. We evaluated it in several simulated data sets, varying in number of samples, mean coverage and population mutation rate, and in two empirical human data sets (N = 8 and N = 63 samples). In our simulations, GBStools improved genotype accuracy more than commonly used filters such as Hardy-Weinberg equilibrium p-values. GBStools is most effective at removing genotype errors in data sets over 100 samples when coverage is 40X or higher, and the improvement is most pronounced in species with high genomic diversity. We also demonstrate the utility of GBS and GBStools for human population genetic inference in Argentine populations and reveal widely varying individual ancestry proportions and an excess of singletons, consistent with recent population growth.

Genomics

The Time-Scale of Recombination Rate Evolution in Great Apes

We present three linkage-disequilibrium (LD)-based recombination maps generated using whole-genome sequencing data of 10 Nigerian chimpanzees, 13 bonobos, and 15 western gorillas, collected as part of the Great Ape Genome Project (Prado-Martinez et al. 2013). Using species-specific PRDM9 sequences to predict potential binding sites, we identified an important role for PRDM9 in predicting recombination rate variation broadly across great apes. Our results are contrary to previous research that PRDM9 is not associated with recombination in western chimpanzees (Auton et al. 2012). Additionally, we show that fewer hotspots are shared among chimpanzee subspecies than within human populations, further narrowing the time-scale of complete hotspot turnover. We quantified the variation in the biased distribution of recombination rates towards recombination hotspots across great apes. We found that correlations between broad-scale recombination rates decline more rapidly than nucleotide divergence between species. We also compared the skew of recombination rates at centromeres and telomeres between species and show a skew from chromosome means extending as far as 10-15 Mb from chromosome ends. Further, we examined broad-scale recombination rate changes near a translocation in gorillas and found minimal differences as compared to other great ape species perhaps because the coordinates relative to the chromosome ends were unaffected. Finally, based on multiple linear regression analysis, we found that various correlates of recombination rate persist throughout primates including repeats, diversity, divergence and local effective population size (Ne). Our study is the first to analyze within-and between-species genome-wide recombination rate variation in several close relatives.

Evolutionary Biology

Illumina TruSeq synthetic long-reads empower de novo assembly and resolve complex, highly repetitive transposable elements

High-throughput DNA sequencing technologies have revolutionized genomic analysis, including the de novo assembly of whole genomes. Nevertheless, assembly of complex genomes remains challenging, in part due to the presence of dispersed repeats which introduce ambiguity during genome reconstruction. Transposable elements (TEs) can be particularly problematic, especially for TE families exhibiting high sequence identity, high copy number, or present in complex genomic arrangements. While TEs strongly affect genome function and evolution, most current de novo assembly approaches cannot resolve long, identical, and abundant families of TEs. Here, we applied a novel Illumina technology called TruSeq synthetic long-reads, which are generated through highly parallel library preparation and local assembly of short read data and achieve lengths of 1.5-18.5 Kbp with an extremely low error rate ([~]0.03% per base). To test the utility of this technology, we sequenced and assembled the genome of the model organism Drosophila melanogaster (reference genome strain y;cn,bw,sp) achieving an N50 contig size of 69.7 Kbp and covering 96.9% of the euchromatic chromosome arms of the current reference genome. TruSeq synthetic long-read technology enables placement of individual TE copies in their proper genomic locations as well as accurate reconstruction of TE sequences. We entirely recovered and accurately placed 4,229 (77.8%) of the 5,434 of annotated transposable elements with perfect identity to the current reference genome. As TEs are ubiquitous features of genomes of many species, TruSeq synthetic long-reads, and likely other methods that generate long reads, offer a powerful approach to improve de novo assemblies of whole genomes.

Genomics