Search bioRxiv⌕ Search

bioRxiv · 10.1101/2025.03.25.645253

Delphy: scalable, near-real-time Bayesian phylogenetics for outbreaks

Abstract

Pathogen genomic analysis is central to tracking, understanding, and containing outbreaks, but complexity and high costs of state-of-the-art (SOTA) phylogenetic tools limit global access and impact. We introduce Delphy, an exact reformulation of Bayesian phylogenetics designed to transform its speed, scalability and accessibility while retaining SOTA accuracy. Delphys central data structure, an Explicit Mutation Annotated Tree, exploits the high sequence similarity in large-scale epidemic datasets for efficient tree exploration and convergence. By reproducing key analyses from recent major epidemics (Ebola, Zika, SARS-CoV-2, mpox, and H5N1), we demonstrate SOTA accuracy with up to 1,000x speedups. Assessing Delphys scalability, we show that a simulated dataset of 100,000 sequences can be analyzed in under a day-the largest such computation to date. We distribute Delphy as a client-side web application, enabling users worldwide to turn raw data into interactive results within minutes, without the data ever leaving the users machine. Delphy automatically identifies key viral lineages and mutations, as well as their emergence and prevalence through time, all with quantified uncertainties derived from a solid theoretical foundation. Delphy shows the power of Bayesian phylogenetics as a fast, accessible frontline tool for tackling future outbreaks.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Varilly, P., Schifferli, M., Yang, K., Burcham, T., Cronan, P., Glennon, O., Jacks, O., Laning, E., Marrs, L., Oba, K., Yeung, S., Parker, E., Omah, I., Pekar, J. E., Luebbert, L., Andersen, K. G., Park, D. J., Schaffner, S. F., MacInnis, B. L., Happi, C., Lemieux, J. E., Ozonoff, A., Mitzenmacher, M. D., Fry, B., Sabeti, P. C.. 2025-03-26. Delphy: scalable, near-real-time Bayesian phylogenetics for outbreaks. https://doi.org/10.1101/2025.03.25.645253

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Interrogation of noncoding schizophrenia risk variants using CRISPR-based functional genomics

Schizophrenia (SCZ) is a highly heritable complex disorder influenced by coding and noncoding genetic variation. Its genetic causes, particularly those involving noncoding variation, are largely unknown. High-throughput CRISPR screens enable dissection of disease-associated loci and identification of noncoding regulatory elements and variants that modulate gene expression. We screened SCZ GWAS loci linked to genes that are also associated in whole-exome sequencing studies to identify regulatory elements and variants impacting expression of disease-relevant genes. We used CRISPRi paired with HCR-FlowFISH to epigenetically silence 333 putative regulatory elements and measure the downstream effects on gene expression of causal SCZ genes, FAM120A, SV2A, and STAG1, in iPSCs and iPSC-derived neurons (iNeurons). We identified 78 regulatory elements that significantly alter expression of a SCZ gene, including noncoding enhancers/silencers as well as promoters of genes and lncRNAs. Pooled prime editing screens interrogated noncoding variant influence on gene expression for SCZ-associated variants and uncharacterized common variants from diverse population studies. We find that a common variant in the promoter of SV2A, rs112851681:A>G (MAF = 3.56%, 1000 Genomes) enhances transcriptional activity in iPSCs and iNeurons. These findings show distinct noncoding mechanisms that map within GWAS signals, and provide a path forward for interrogating noncoding regulatory elements and variants in disease loci.

genomics↗

The automated eukaryotic pangenome pipeline EukPan reveals accessory genome differentiation beyond core-gene phylogeny in Aspergillus oryzae

Pangenome analysis reveals recurrent gene-content variation beyond a single reference genome, but its application to eukaryotes is constrained by inconsistent gene annotation. ANNEVO predicts gene models from genome FASTA assemblies without RNA-seq data. We developed EukPan, an automated post-annotation pipeline that standardizes GFF/GTF files, selects representative isoforms, constructs proteomes, infers orthogroups, builds a concatenated single-copy core-protein alignment, and summarizes shared accessory orthogroups while excluding orthogroups detected in only one genome. Applied with ANNEVO to 123 Aspergillus oryzae genomes, EukPan identified 11,245 core and 4,407 shared accessory orthogroups. The core-protein phylogeny broadly recovered the reported A-H classification, whereas accessory-genome analyses clearly separated the 33 group-A strains from the other 90 strains. Directional analysis identified 62 group-A-associated and 158 group-A-depleted orthogroups, with major facilitator superfamily (MFS) transporter and fungal Zn2Cys6 transcription-factor domains prominent in the depleted set. Among 93 orthogroups present in all non-A strains and absent from all group-A strains, 59 mapped to 10 segments of RIB40, the standard A. oryzae reference genome and a non-A (group-F) strain. EukPan therefore enables reproducible, coordinated core- and accessory-pangenome analysis from eukaryotic genome assemblies.

genomics↗

Multi-Omics analysis provides crucial insights into ecological adaptation to dryland of a dominant grass (Psammochloa villosa, Poaceae) in Northwest China

Desertification exerts dramatic selection pressures on the evolution of plants. Despite the key role of ecological adaptation by natural selection to arid grasslands and subsequent intraspecific divergence, specific mechanisms driving this process remain poorly understood. Psammochloa villosa, a perennial forage grass endemic to the arid grasslands in Northwest China, where it thrives in shifting and semi-fixed sand land due to its exceptional drought tolerance, provides an ideal system to study adaptive evolution to aridity. In our study, we assembled a high-quality, chromosome-scale genome and conducted genomic resequencing of 42 populations across its major distribution. The genome assembly, which is approximately 1.55 Gb in size, has a super-scaffold N50 of 66.79 Mb, with 75.84% of the sequences identified as transposable elements. Coalescent phylogeny and genomic collinearity analyses strongly supported that P. villosa and Neotrinia splendens, as the closest taxa, shared a recent whole-genome duplication (WGD) event occurring approximately 18-20 Mya and followed by their divergence around ~11.2 Mya. Based on ancestral grass karyotype (AGK) reconstruction from synteny analysis, our results suggest that, relative to the AGK after the {rho}-WGD event, P. villosa and N. splendens underwent similar chromosomal restructuring and lineage-specific retention of numerous copies, providing a genomic basis for potential ecological adaptation and intraspecific diversification. The expanded XTH family, which encodes enzymes mediating xyloglucan endotransglucosylation and hydrolysis and thereby regulating xyloglucan remodeling, showed strong transcriptional responses under PEG-6000 treatment, suggesting that retained copies may be associated with xerophytic adaptation in P. villosa. Together, these findings suggest that WGD-derived gene retention created a delayed reservoir of genetic diversity that was later shaped by desertification and Qinghai-Xizang Plateau environmental changes, contributing to climate-associated genomic islands and intraspecific differentiation.

genomics↗