Search bioRxivSearch

Biology subjects

Hickey, J.

Publications and source records attributed to Hickey, J..

3 recordsLinked to original sources

Impact of index hopping and bias towards the reference allele on accuracy of genotype calls from low-coverage sequencing

BackgroundInherent sources of error and bias that affect the quality of the sequence data include index hopping and bias towards the reference allele. The impact of these artefacts is likely greater for low-coverage data than for high-coverage data because low-coverage data has scant information and standard tools for processing sequence data were designed for high-coverage data. With the proliferation of cost-effective low-coverage sequencing there is a need to understand the impact of these errors and bias on resulting genotype calls.\n\nResultsWe used a dataset of 26 pigs sequenced both at 2x with multiplexing and at 30x without multiplexing to show that index hopping and bias towards the reference allele due to alignment had little impact on genotype calls. However, pruning of alternative haplotypes supported by a number of reads below a predefined threshold, a default and desired step for removing potential sequencing errors in high-coverage data, introduced an unexpected bias towards the reference allele when applied to low-coverage data. This bias reduced best-guess genotype concordance of low-coverage sequence data by 19.0 absolute percentage points.\n\nConclusionsWe propose a simple pipeline to correct this bias and we recommend that users of low-coverage sequencing be wary of unexpected biases produced by tools designed for high-coverage sequencing.

bioinformatics

Parentage assignment with low density array data and low coverage sequence data

In this paper we evaluate using genotype-by-sequencing (GBS) data to perform parentage assignment in lieu of traditional array data. The use of GBS data raises two issues: First, for low-coverage GBS data, it may not be possible to call the genotype at many loci, a critical first step for detecting opposing homozygous markers. Second, the amount of sequencing coverage may vary across individuals, making it challenging to directly compare the likelihood scores between putative parents. To address these issues we extend the probabilistic framework of Huisman (2017) and evaluate putative parents by comparing their (potentially noisy) genotypes to a series of proposal distributions. These distributions describe the expected genotype probabilities for the relatives of an individual. We assign putative parents as a parent if they are classified as a parent (as opposed to e.g., an unrelated individual), and if the assignment score passes a threshold. We evaluated this method on simulated data and found that (1) high-coverage GBS data performs similarly to array data and requires only a small number of markers to correctly assign parents and (2) low-coverage GBS data (as low as 0.1x) can also be used, provided that it is obtained across a large number of markers. When analysing the low-coverage GBS data, we also found a high number of false positives if the true parent is not contained within the list of candidate parents, but that this false positive rate can be greatly reduced by hand tuning the assignment threshold. We provide this parentage assignment method as a standalone program called AlphaAssign.

genetics

A strategy to exploit surrogate sire technology in livestock breeding programs

In this work, we performed simulations to develop and test a strategy for exploiting surrogate sire technology in animal breeding programs. Surrogate sire technology allows the creation of males that lack their own germline cells, but have transplanted spermatogonial stem cells from donor males. With this technology, a single elite male donor could give rise to huge numbers of progeny, potentially as much as all the production animals in a particular time period.\n\nOne hundred replicates of various scenarios were performed. Scenarios followed a common overall structure but differed in the strategy used to identify elite donors and how these donors were used in the product development part.\n\nThe results of this study showed that using surrogate sire technology would significantly increase the genetic merit of commercial sires, by as much as 6.5 to 9.2 years worth of genetic gain compared to a conventional breeding program. The simulations suggested that a strategy involving three stages (an initial genomic test followed by two subsequent progeny tests) was the most effective of all the strategies tested.\n\nThe use of one or a handful of elite donors to generate the production animals would be very different to current practice. While the results demonstrate the great potential of surrogate sire technology there are considerable risks but also other opportunities. Practical implementation of surrogate sire technology would need to account for these.

genetics