Search bioRxiv⌕ Search

Biology subjects

Tallman, S.

Publications and source records attributed to Tallman, S..

2 recordsLinked to original sources

Analysis-ready VCF at Biobank scale using Zarr

BackgroundVariant Call Format (VCF) is the standard file format for interchanging genetic variation data and associated quality control metrics. The usual row-wise encoding of the VCF data model (either as text or packed binary) emphasises efficient retrieval of all data for a given variant, but accessing data on a field or sample basis is inefficient. Biobank scale datasets currently available consist of hundreds of thousands of whole genomes and hundreds of terabytes of compressed VCF. Row-wise data storage is fundamentally unsuitable and a more scalable approach is needed. ResultsZarr is a format for storing multi-dimensional data that is widely used across the sciences, and is ideally suited to massively parallel processing. We present the VCF Zarr specification, an encoding of the VCF data model using Zarr, along with fundamental software infrastructure for efficient and reliable conversion at scale. We show how this format is far more efficient than standard VCF based approaches, and competitive with specialised methods for storing genotype data in terms of compression ratios and single-threaded calculation performance. We present case studies on subsets of three large human datasets (Genomics England: n=78,195; Our Future Health: n=651,050; All of Us: n=245,394) along with whole genome datasets for Norway Spruce (n=1,063) and SARS-CoV-2 (n=4,484,157). We demonstrate the potential for VCF Zarr to enable a new generation of high-performance and cost-effective applications via illustrative examples using cloud computing and GPUs. ConclusionsLarge row-encoded VCF files are a major bottleneck for current research, and storing and processing these files incurs a substantial cost. The VCF Zarr specification, building on widely-used, open-source technologies has the potential to greatly reduce these costs, and may enable a diverse ecosystem of next-generation tools for analysing genetic variation data directly from cloud-based object stores, while maintaining compatibility with existing file-oriented workflows. Key PointsO_LIVCF is widely supported, and the underlying data model entrenched in bioinformatics pipelines. C_LIO_LIThe standard row-wise encoding as text (or binary) is inherently inefficient for large-scale data processing. C_LIO_LIThe Zarr format provides an efficient solution, by encoding fields in the VCF separately in chunk-compressed binary format. C_LI

bioinformatics↗

Long-term evolution of Streptococcus mitis and Streptococcus pneumoniae leads to higher genetic diversity within rather than between human populations

Evaluation of the apportionment of genetic diversity of bacterial commensals within and between populations is an important step in the characterization of their evolutionary potential. Recent studies showed a correlation between the genomic diversity of human commensal strains and that of their host, but the strength of this correlation and of the geographic structure among populations is a matter of debate. Here, we studied the genomic diversity and evolution of the phylogenetically related oronasopharingeal healthy-carriage Streptococcus mitis and Streptococcus pneumoniae, whose lifestyles range from stricter commensalism to high pathogenic potential. A total of 119 S. mitis genomes showed higher within- and among-host variation than 810 S. pneumoniae genomes in European, East Asian and African populations. Summary statistics of the site-frequency spectrum for synonymous and non-synonymous variation and ABC modelling showed this difference to be due to higher historical population effective size (Ne) in S. mitis, whose genomic variation been maintained close to mutation-drift equilibrium across (at least many) generations, whereas S. pneumoniae has been expanding from a smaller ancestral population and has been subjected to adaptive selection. Strikingly, both species show limited differentiation among populations. As genetic differentiation is inversely proportional to the product of effective population size and migration rate (Nem), we argue that large Ne have led to similar differentiation patterns, even if m is very low for S. mitis. We conclude that more diversity within than among human populations and limited population differentiation must be common features of the human microbiome due to large Ne.

evolutionary biology↗