bioRxiv · 10.1101/2025.02.05.636549
IGD: A simple, efficient genotype data format
Abstract
MotivationWhile there are a variety of file formats for storing reference-sequence-aligned genotype data, many are complex or inefficient. Programming language support for such formats is often limited. A file format that is simple to understand and implement - yet fast and small - is helpful for research on highly scalable bioinformatics. ResultsWe present the Indexable Genotype Data (IGD) file format, a simple uncompressed binary format that can be more than 100 times faster and 3.5 times smaller than vcf.gz on Biobank-scale whole-genome sequence data. The implementation for reading and writing IGD in Python is under 350 lines of code, which reflects the simplicity of the format. AvailabilityA C++ library reading and writing IGD, and tooling to convert .vcf.gz files, can be found at https://github.com/aprilweilab/picovcf. A Python library is at https://github.com/aprilweilab/pyigd
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
DeHaas, D., Wei, X.. 2025-02-08. IGD: A simple, efficient genotype data format. https://doi.org/10.1101/2025.02.05.636549
Cite the original work for its findings. Save a collection to share your selection of sources.