Pervasive contaminations in sequencing experiments are a major source of false genetic variability: a Mycobacterium tuberculosis meta-analysis
Contaminant DNA is a well-known confounding factor in molecular biology and in genomic repositories. Strikingly, analysis workflows for whole-genome sequencing (WGS) data usually neglect the errors introduced by potential contaminations. We performed a comprehensive evaluation of the extent and impact of contaminant DNA in WGS by analyzing more than 4,000 bacterial samples from 20 different studies. We found that contaminations are pervasive and can introduce large biases in variant analysis. We showed that these biases can translate in hundreds of false positive and negative SNPs, even for samples with slight contaminations. Studies investigating complex biological traits from sequencing data can be completely biased if contaminations are neglected during the bioinformatic analysis. We used both real and simulated data to evaluate and implement reliable, contamination-aware analysis pipelines. Our results urge for the implementation of such pipelines as sequencing technologies consolidate as a precision tool in the research and clinical context.