Search bioRxiv⌕ Search

Biology subjects

Denni, S.

Publications and source records attributed to Denni, S..

2 recordsLinked to original sources

Simulating population pangenomes under coalescent demographic models with MSpangenome

MotivationPangenome variation graphs (PVGs) are increasingly used to represent genomic diversity, yet there is currently no general framework for generating population pangenomes directly from explicit evolutionary histories. Existing simulators typically focus on individual classes of variation and do not integrate these variations within a genealogy-aware framework driven by explicit demographic histories. As a result, evaluating pangenome methods in realistic population-genetic settings remains challenging, and benchmark datasets with known evolutionary ground truth are scarce. ResultsWe present MSpangenome, a genealogy-aware framework that bridges coalescent population genetic simulations and pangenome graph analyses. The pipeline combines ancestry simulation with msprime and a de novo graph construction algorithm to generate PVGs directly from simulated genealogies. By explicitly modeling recombination, demographic history and incomplete lineage sorting, MSpangenome produces structurally complex pangenomes in which nested and overlapping structural variants emerge naturally from the underlying genealogies, while their evolutionary history and graph topology remain known by construction. This provides a general framework for generating realistic population pangenomes and establishing ground-truth datasets for methodological evaluation. We demonstrate its utility by generating population-scale pangenomes and using them as controlled references to benchmark the widely used graph construction tools, PGGB and Minigraph-Cactus. Our analyses reveal contrasting performance regimes across levels of sequence diversity, sample sizes and classes of structural variation, highlighting the value of simulation-based benchmarking for identifying reconstruction errors that are hard to detect using empirical datasets alone. Availability and implementationMSpangenome is implemented in Python, fully containerized, freely available at https://forge.inrae.fr/pangepop/MSpangepop and mirrored at https://github.com/inrae/MSpangepop.

bioinformatics↗

Building pangenomes for domesticated and wild tree species: genomic complexity and strategies

Long-read sequencing and pangenomics are revolutionizing crop research by providing more complete genome information and revealing crucial structural variations linked to important agricultural traits. Building on recent advances in intraspecific pangenome construction, this study addresses the challenge of creating broader, cross-taxon pangenomes, using the Armeniaca taxonomic section as a model. Leveraging a diverse panel of genome assemblies, we constructed a pangenome graph and cataloged associated single nucleotide polymorphisms (SNPs) and structural variants. We characterized the diversity of these variants and assessed the extent to which different taxa contribute to overall pangenome expansion. Additionally, we evaluated the performance of low-depth sample mapping to the graph-based reference, highlighting key technical limitations that may affect the quality of downstream analyses. We further identified specific subsets of SVs that exhibit associations with particular classes of transposable elements. As a case study illustrating the potential functional and phenotypic relevance of graph-derived SVs, we examined the genomic configuration of the DAM locus within the Armeniaca pangenome.

bioinformatics↗