bioRxiv · 10.1101/518308
Cleaning genotype data from Diversity Outbred mice
Abstract
Data cleaning is an important first step in most statistical analyses, including efforts to map the genetic loci that contribute to variation in complex quantitative traits. Data cleaning is more difficult in multiparent populations (experimental crosses derived from more than two founder strains), as individual SNP markers have incomplete information. We describe our process for cleaning SNP microarray genotype data from Diversity Outbred (DO) mice, illustrated with genotype data from the MegaMUGA array on a set of 291 DO mice. We first consider the proportion of missing genotypes in each mouse, as an indicator of sample quality. We use microarray probe intensities for SNPs on the X and Y chromosomes to check mouse sex. We use the proportion of matching SNP genotypes in pairs of mice to detect sample duplicates. We use a hidden Markov model to calculate genotype probabilities across the genome for each mouse, which we use to estimate the number of crossovers in each mouse, and to identify potential genotyping errors. The estimated genotyping error rate, by mouse sample and by SNP marker, is a useful diagnostic. We also consider the SNP genotype frequencies by mouse and by marker, with markers grouped according to the minor allele frequency in the founder strains. For markers with high apparent error rates, a scatterplot of the allele intensities from the microarray, colored by both observed and predicted genotypes, can be revealing about the underlying cause. But the detection of low-quality samples is more important than the detection of low-quality markers.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Broman, K. W., Gatti, D. M., Svenson, K. L., Sen, S., Churchill, G. A.. 2019-01-11. Cleaning genotype data from Diversity Outbred mice. https://doi.org/10.1101/518308
Cite the original work for its findings. Save a collection to share your selection of sources.