bioRxiv · 10.1101/060822
FastGT: from raw sequence reads to 30 million genotypes in less than an hour
Abstract
We have developed a computational method that counts the frequencies of unique k-mers in FASTQ-formatted genome data and uses this information to infer the genotypes of known variants. FastGT can detect the variants in a 30x genome in less than 1 hour using ordinary low-cost server hardware. The overall concordance with the genotypes of two Illumina \"Platinum\" genomes1 is 99.96%, and the concordance with the genotypes of the Illumina HumanOmniExpress is 99.82%. Our method provides k-mer database that can be used for the simultaneous genotyping of approximately 30 million single nucleotide variants (SNVs), including >23,000 SNVs from Y chromosome. The source code of FastGT software is available at GitHub (https://github.com/bioinfo-ut/GenomeTester4/).
Source connections
Explore related subjects
Keep this discovery
Fanny-Dhelia Pajuste, Lauris Kaplinski, Märt Möls, Tarmo Puurand, Maarja Lepamets, Maido Remm. 2016-06-28. FastGT: from raw sequence reads to 30 million genotypes in less than an hour. https://doi.org/10.1101/060822
Cite the original work for its findings. Save a collection to share your selection of sources.