Search bioRxiv⌕ Search

bioRxiv · 10.1101/2025.05.28.656711

Benchmarking Genomic Variant Calling Tools in Inbred Mouse Strains: Recommendations and Considerations

Abstract

With the growing affordability of whole genome sequencing, variant identification has become an increasingly common task, but there are many challenges due to both technical and biological factors. In recent years, the number of software packages available for variant calling has rapidly increased. Understanding the benefits and drawbacks of different tools is important in setting leading practices and highlighting limitations. These considerations are crucial in model organism research, as many variant calling programs assume outbred genomes and implicit heterozygosity, which may not apply to inbred laboratory models. Here, we present an analysis of variant calling tools and their performance in the simulated genomes of the C57BL/6J inbred laboratory mouse and nine non-reference laboratory strains. Our findings reveal a tradeoff between the recall and precision of tools. Balancing these considerations, we show that an optimal call set is obtained by using an ensemble approach, but specific variant calling recommendations vary by strain and analytical goals. Further, we highlight filters improving the performance of different variant calling tools, both for the discovery of rare variants and in the discovery of strain polymorphisms. In summary, our work provides best practices for calling and filtering genomic variants in inbred organisms, particularly laboratory mice. Article SummaryIdentifying mutations and rare genetic variants is a central task for modern genomics. Many computational tools exist for variant detection, but their performance varies across diverse applications. Further, few variant calling tools have been benchmarked against inbred genomes, which are commonly used for research. To address this, we evaluated five variant calling tools using simulated data from ten diverse inbred mouse strains. We show that variant detection, recall, and precision vary across tools and mouse strains, and that an ensemble approach improves confidence in detected mutations. Our findings offer a set of best practices for variant calling in inbred organisms across diverse analytical applications.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Garretson, A. C., Blanco-Berdugo, L., Roberts, A., Dumont, B. L.. 2025-05-31. Benchmarking Genomic Variant Calling Tools in Inbred Mouse Strains: Recommendations and Considerations. https://doi.org/10.1101/2025.05.28.656711

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

spatialMET: an open and scalable framework for spatial metabolomics analysis

Mass spectrometry imaging (MSI) enables spatially resolved metabolomics in intact tissue sections, but analysis remains challenging at scale. Existing MSI workflows often require users to combine multiple software tools, while others rely on proprietary vendor software that limits interoperability and reproducibility. To address these challenges, we developed spatialMET, an open-source framework that provides an end-to-end workflow for MSI analysis. spatialMET provides a unified platform for preprocessing, spatial domain detection, and visualization. Downstream analyses include differential abundance testing, spatial autocorrelation and gradient analysis, dimensionality reduction, and correlation network analysis. Spatial domain detection uses hcdist, a C-based hierarchical clustering implementation that substantially reduces runtime and memory use relative to existing R-based approaches. spatialMET can be run through an interactive R Shiny application or as a standalone command-line workflow for larger datasets or high-performance computing environments. Applied to mouse small cell lung cancer MALDI-MSI data containing 284,673 pixels, spatialMET identified tumor-associated, stromal, and adjacent lung spatial domains that aligned with matched histology. Differential abundance analysis identified 117 m/z features that differed between tumor and stromal regions, while spatial autocorrelation analyses revealed spatially structured abundance patterns. Applying spatialMET to mouse lung adenocarcinoma data from an entire lung lobe containing 338,477 pixels further demonstrated scalability and captured spatial heterogeneity across tumor and surrounding lung tissue. In summary, spatialMET provides a scalable, open-source framework for end-to-end spatial metabolomics analysis, and it is distributed as a Docker container for reproducible deployment. Source code and installation instructions are available at https://github.com/biodatalab/spatialMET.

bioinformatics↗

Probing the transcriptome response to shivering in skeletal muscle using a multilayered bioinformatics approach

Cold acclimation holds therapeutic potential for improving metabolic health. We previously demonstrated that repeated cold-induced shivering enhances insulin sensitivity in humans. However, the molecular pathways that underlie the skeletal muscle shivering response, and how these relate to beneficial physiological effects, remain poorly understood. In this study, we combined complementary bioinformatics approaches to allow in-depth analysis of the transcriptomic response of human skeletal muscle to repeated shivering. We identified a robust transcriptional signature and show a sex-specific component in the shivering skeletal muscle response, which seemed to diminish following cold adaptation. Our findings provide mechanistic insights into cold-induced muscle adaptations, shed light on potential interesting molecular targets for further investigation, and emphasize the importance of including both sexes in future cold acclimation studies.

bioinformatics↗

An Information Geometry approach to model topological trajectories and Gene Expression Radius from UMAP geometry.

Understanding the relationship between gene expression dynamics and cellular identity remains a central challenge in single cell biology. Here, we introduce a novel computational and mathematical framework that integrates information geometry, fuzzy topology, and UMAP analysis to model gene expression landscapes derived from single cell RNA sequencing data. We formalize gene expression data as a fuzzy topological space, where interactions between expression points are governed by probabilistic distributions inspired by manifold learning approaches such as UMAP. Within this framework, we define an information geometric structure through a Fisher metric induced by these distributions, enabling the computation of geodesic trajectories that capture cellular differentiation processes. A key contribution of this work is the derivation of analytical conditions, expressed as expression radius formulas, that characterize local neighborhoods in gene expression space. These conditions allow for the identification of genes associated with stem cell states and predictions in transitional cell types in future work. Application of the proposed framework to single cell datasets reveals biologically meaningful gene sets enriched in key regulatory pathways and transcription factors, demonstrating the capacity of our approach to uncover latent structure in complex gene expression data. Our results suggest that integrating differential geometry with statistical learning theory offers a powerful paradigm for modeling genotype and phenotype relationships and cellular state transitions, with potential implications for precision medicine and systems biology.

bioinformatics↗