Search bioRxiv⌕ Search

bioRxiv · 10.1101/2024.12.05.626966

Gfa2bin enables graph-based GWAS by converting genome graphs to pan-genomic genotypes

Abstract

Variation graphs offer superior representation of genomic diversity compared to traditional linear reference genomes, capturing complex features that are otherwise inaccessible to analysis. It seems self-evident that integrating these graphs with genome-wide association studies (GWAS) should enable more comprehensive understanding of genetic landscapes, potentially uncovering novel associations between genetic variations and traits. This approach takes full advantage of rich genomic information, thereby providing deeper insights into the genetic base of complex traits. Our tool, gfa2bin, offers multiple methods to (i) genotype variation graphs and (ii) convert the genotypes to well-established data formats for genome-wide association studies (GWAS). We demonstrate that variation graphs are feasible alternatives to traditional linear references for GWAS. Our case study using Arabidopsis thaliana and 1,695 traits shows that our approach complements SNP-based approaches, often identifying additional associations, with all associations having on average higher significance compared to SNP-based approaches. gfa2bin is implemented in Rust. Commented source code is available under MIT license at https://github.com/MoinSebi/gfa2bin. Examples of how to run gfa2bin are provided in the documentation. We added several Python scripts and a Snakemake pipeline for easy processing of our tool using larger data sets. In addition, we recommend using packing (https://github.com/MoinSebi/packing) for reduced storage and preprocessing (normalization) of sequence-to-graph alignments coverage.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Vorbrugg, S., Bezrukov, I., Bao, Z., Xian, W., Weigel, D.. 2024-12-09. Gfa2bin enables graph-based GWAS by converting genome graphs to pan-genomic genotypes. https://doi.org/10.1101/2024.12.05.626966

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Integrating Genomic Annotations and Traits Dependencies for single-nucleotide polymorphisms Prioritization with Causal Concept Bottleneck Models

Predicting common traits from single-nucleotide polymorphism (SNPs) data is challenging due to polygenicity, small effect sizes, and the presence of potentially mediated or spurious cross-trait associations. We propose a modeling approach that combines genomic annotations with known cross-trait relations by leveraging Causally Reliable Concept Bottleneck Models (C2BM), a deep learning architecture that factors the joint trait distribution over a graph of interpretable concepts. This design allows trait predictions to leverage information from other observed traits in addition to genomic inputs. Furthermore, the interpretable architecture of the model enables us to investigate how specific trait-trait relationships influence SNP-level predictions. We evaluate the approach on a multi-trait GWAS dataset covering five traits and show that C2BM improves predictions when ground-truth labels for related traits are available. Moreover, by analyzing variations in how trait-trait relationships influence predictions, we postulate that such differences may reflect the presence or absence of shared genetic mechanisms or indirect effects. Accepted at the CIBB 2026 conference (https://cibb2026.teralab.ai/)

bioinformatics↗

Cophylogeny simulators are not interchangeable: similarities, differences and structural biases in synthetic host-symbiont

Synthetic data are becoming increasingly important for computational studies of cophylogeny, including machine learning inference, benchmarking, and method testing. Several generators have been proposed to produce such data, but each relies on different assumptions about host-symbiont coevolution. These assumptions are often implicit and rarely examined, even though results can depend strongly on the synthetic model being used. In this article, we present a systematic structural analysis of representative cophylogeny generators under controlled scenarios. The goal is to make their assumptions explicit and to understand how these choices shape the synthetic data they produce as well as the conclusions that may be drawn from them.

bioinformatics↗

Evolutionary Diversification of Nitric Oxide Signaling Components Across Metazoa: A Comparative Phylogenomic Analysis

Nitric oxide (NO) is an evolutionarily ancient gaseous signaling molecule in animals, yet the evolutionary history of multiple components spanning NO synthesis, substrate regulation, sensing, and signal termination has not been examined in a single integrated phylogenetic framework across Metazoa. Here, we trace the phylogenetic and gene-tree/species-tree histories of ten core NO-pathway components across 89 eukaryotic proteomes spanning Amoebozoa, Excavata, Fungi, Archaeplastida, and Opisthokonta - including Porifera, Placozoa, Cnidaria, Ctenophora, and Bilateria - with a focus on the 69 opisthokont species that anchor the animal comparisons. The results reveal a strikingly modular evolutionary architecture. Nitric oxide synthase (NOS) is broadly conserved across bilaterian lineages, and reconciliation analyses show that NOS diversification was driven predominantly by speciation rather than lineage-specific duplication - supporting the relative conservation of NOS across the sampled bilaterian lineages . By contrast, the arginine-recycling enzymes ASS1 and ASL show ancestral duplications and inferred secondary losses in specific lineages, while arginase isoforms (ARG1/ARG2) and the cGMP-degrading enzyme PDE5A underwent extensive, independent duplications across metazoan groups. GUCY1A1- and GUCY1B1-like sequences were recovered across several metazoan lineages, but the two subunits exhibited partially divergent evolutionary trajectories. Together, these patterns support a model in which a comparatively conserved NO-producing component coexists with more dynamic diversification of associated pathway gene families, providing an evolutionary framework for investigating how NO-cGMP signaling may have been differentially deployed in nervous systems.

bioinformatics↗