Search bioRxiv⌕ Search

bioRxiv · 10.1101/2025.04.17.649315

μSeq: Universal mutation rate quantification via deep sequencing of a single clonal expansion.

Abstract

Understanding and quantifying mutational processes is fundamental for studying evolution in both clinical and experimental contexts. However, current methods are labor intensive and often lack robustness, particularly in mammalian cells. For example, current approaches that rely on subclonal mutations derived directly from patient samples often result in inconsistent or biased outcomes, leading to potentially unreliable conclusions on adaptive dynamics. We used patient-derived colorectal cancer organoids to introduce {micro}Seq, a universal framework for inferring mutation rates in diverse biological systems. Our approach extracts mutation rates from deep sequencing of single clonal expansions, with a time gain of ten-fold or more compared to a mutation accumulation line, at the cost of three billion read whole genome sequencing. {micro}Seq relies on four critical components: (i) a controlled experimental setup enabling validation (ii) a quantitative estimate inspired by the classic Luria-Delbruck spectrum for subclonal mutations derived from population dynamics, (iii) robust statistical models accounting for sampling noise and sequencing errors, and (iv) a data analysis pipeline for subclonal mutation detection that compares endpoint populations to a closely related ancestor. Building on the legacy of Luria and Delbruck, our model adopts core concepts from statistical physics--stochasticity, fluctuation spectra, and inference under noise--to construct a rigorous and scalable inference framework. We demonstrate that the Luria-Delbruck distribution extends to subclonal mutations in expanding populations, and show how this generalization enables robust estimation of the underlying mutation rate. Our models establish precise requirements for sequencing depth, genome size, and mutation frequency detection necessary for accurate mutation rate estimates. Crucially, we show that failing to meet these criteria can lead to errors spanning several orders of magnitude suggesting that biases in patients data arise from analyzing incorrect frequency intervals and from lack of a closely related reference. We validate our approach using parallel mutation accumulation experiments in colorectal cancer organoids, finding mutation rate estimates consistent with previous studies. However, we find that mutation accumulation lines operate under purifying selection in both yeast and human organoids, contradicting the long standing assumption of neutrality in such experiments. This insight has important implications for both evolutionary biology and cancer evolution. Finally, to demonstrate the adaptability of {micro}Seq, we apply it to yeast, leveraging multiple independent replicates to compensate for its much smaller genome, as well as in mouse xenografts, which feature much more complex in vivo population dynamics. The robustness and broad applicability of {micro}Seq establish it as a powerful and universal tool for mutation rate quantification, and imply that existing claims based on subclonal mutations from patient samples must be revisited. Short AbstractUnderstanding and quantifying mutational processes is fundamental to studying evolution in clinical and experimental contexts. However, current methods are labor intensive and often lack robustness, particularly in mammalian cells. For example, quantification of subclonal mutations directly from patient samples produces inconsistent estimates, leading to unreliable conclusions about adaptive dynamics. Using patient derived colorectal cancer organoids as a model system, we developed {micro}Seq, a novel framework to estimate mutation rates from deep sequencing of a single clonal expansion with a close reference. {micro}Seq builds on the Luria Delbruck model, combining statistical physics concepts with modern sequencing and inference techniques. It integrates (i) a controlled experimental setup, (ii) robust statistical and population dynamics models, and (iii) a data analysis pipeline for subclonal mutation detection. We establish precise requirements for sequencing depth, genome size, and mutation frequency to ensure accuracy despite sequencing errors. {micro}Seq enables accurate estimation of mutation rates as low as 10-9 mutations per base pair per generation, and we validate it across species using yeast data. Critically, {micro}Seq reveals that mutation accumulation lines undergo purifying selection in both human organoids and yeast, It underscores the need to reconsider claims based on subclonal mutations in patient samples.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Pompei, S., Geroldi, A., Rivetti, P., Grassi, E., Vurchio, V., Tallarico, G., Corti, G., Tattini, L., Liti, G., Bertotti, A., Cosentino Lagomarsino, M.. 2025-04-22. μSeq: Universal mutation rate quantification via deep sequencing of a single clonal expansion.. https://doi.org/10.1101/2025.04.17.649315

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Geometry of antigenic evolution improves influenza vaccine selection

Anticipating antigenic evolution is essential for selecting effective seasonal influenza A/H3N2 vaccine strains. To this end, we integrated hemagglutination-inhibition and neutralization titers spanning 2002 to 2025 into a unified Bayesian antigenic map. The map resolves twelve antigenic clusters advancing in discrete steps, with several clusters co-circulating in most seasons. In 15 of 21 seasons, the WHO-recommended vaccine belonged to an earlier cluster than the dominant circulating cluster. The direction of each vaccine update relative to recent viral drift predicted vaccine effectiveness one season ahead in out-of-sample forecasts. Antigenic distance, the conventional measure of vaccine-virus match, was weakly associated with effectiveness until update direction was accounted for. Retrospectively ranking candidate strains by predicted effectiveness would have selected a strain predicted to outperform the WHO recommendation in every season, raising mean predicted effectiveness by 10 percentage points.

evolutionary biology↗

Evolutionary replay of duplicate-gene retention across independent whole-genome duplications

Whole-genome duplications repeatedly expose ancestral gene lineages to the same broad evolutionary outcome-retention or loss of duplicated copies-but it remains unclear whether this history replays similarly across evolutionary scales. We placed duplicate retention in shared hierarchical orthologous-group coordinates and compared percentile ranks defined within each event-wide mapped universe. Three independent angiosperm whole-genome duplications showed reproducible replay (global rank effect T-replay = 0.210, bootstrap 95% confidence interval 0.172-0.248; permutation P = 1/100,001). A plant reference-panel score specified before target outcomes were examined predicted retention after the Apple/Pear duplication ({rho} = 0.169, n = 373). Deep transfer was heterogeneous: the teleost-genome-duplication estimate was positive but unresolved ({rho} = 0.107, n = 151, 95% confidence interval -0.050 to 0.260), whereas transfer to the ancient budding-yeast whole-genome duplication (yeast WGD) was supported ({rho} = 0.280, n = 186). Independently reconstructed animal outcomes also replayed between teleost and Stylommatophora duplications (r = 0.226, n = 146, P = 0.00326), although the effect remained below a prespecified strong-effect threshold. A strict plant-animal comparison was limited to 25 deeply one-to-one lineages and was unresolved (r = 0.033, 95% confidence interval -0.303 to 0.340). Thus, ancestral gene-lineage identity contributes reproducibly to duplicate retention after independent whole-genome duplications, but replay is structured by evolutionary lineage and modified by event-specific history rather than governed by one universal gene-fate ranking.

evolutionary biology↗

A Hymenoptera-restricted gene mediating ant castes co-opts deeply conserved machinery to control organ size

Lineage-specific genes are widespread and have been implicated as phenotypic innovation inducers, but how they acquire complex developmental functions remains poorly understood. Ant queens and workers develop dramatically different organ sizes from identical genomes under juvenile hormone (JH) control, yet the molecular effectors translating JH signalling into caste-specific organ growth remain unknown. Here we identify torch, a Hymenoptera-restricted gene, as the most consistently gyne-biased and JH-responsive gene across 68 ant species. Knockdown of torch in virgin queens of Monomorium pharaonis produces a worker-like, multi-organ growth-restricted phenotype. Mechanistically, torch harbours an E-box-like motif activated by the JH receptor Gce-Tai and acts as a GA-repeat-binding transcription factor that regulates Hippo signalling, the deeply conserved organ-size control pathway in animals. Expressing torch heterologously in mice and a growth-restricted Drosophila background shows that the gene retained its general growth-promoting activity across more than 700 million years of animal evolution in lineages that lack the gene, establishing that its function is mediated through conserved rather than ant-specific machinery. A lineage-specific gene can therefore acquire complex morphogenetic function by co-opting ancient organ-size circuitry, providing a general route by which novel genes can drive phenotypic innovation.

evolutionary biology↗