Search bioRxiv⌕ Search

Biology subjects

Santacatterina, G.

Publications and source records attributed to Santacatterina, G..

4 recordsLinked to original sources

Scalable, fast and accurate differential gene expression testing from millions of cells of multiple patients

Since the development of DNA microarrays and later RNA bulk sequencing, testing with statistically independent samples has been the standard method for detecting genes with different transcription patterns. Single-cell assays challenge these assumptions because individual cells are statistically dependent, and all proposed methodologies present mathematical limitations or computational bottlenecks that prevent a seamless integration of data from many cells and patients simultaneously. In this work, we solve this crucial limitation by introducing a Bayesian framework that retrieves the independence structure at the level of individual patients, separating differences across individuals from actual transcriptional differences. Leveraging multi-GPU and variational inference, our approach excels across different experimental designs and scales to analyse over 10 million cells. This framework enables single-cell differential expression analysis that can finally integrate datasets from large clinical cohorts, atlas projects, or drug-response screens with thousands of samples and millions of cells.

bioinformatics↗

A reference-free strategy for circulating tumor DNA detection from whole-genome sequencing data

Circulating tumor DNA (ctDNA) is emerging as a promising biomarker for postoperative monitoring of cancer patients. Precise estimation of circulating tumor fraction is crucial for evaluating treatment effects and timely detection of disease recurrence. All current ctDNA detection methods that utilize whole-genome sequencing (WGS) data rely on the reference genome alignment of sequencing reads and often apply separate tools for detecting different variant types. However, various bioinformatic analysis confounders and the application of external variant calling tools could be avoided by analyzing k-mers from unaligned sequencing reads. While k-mer-based methods have successfully been applied for somatic variant validation and detection, the potential of k-mer-based ctDNA detection is unexplored. We have developed a tumor-informed reference-free ctDNA detection tool called ctDNAmer that detects tumor-specific somatic variation directly from unaligned sequencing data by identifying k-mers unique to the tumor DNA. ctDNAmer detects variant information across the genome by comparing the primary tumor and germline WGS data and accounts for sample-specific germline variability and technical noise in the same framework. We tested the utility of ctDNAmer for tumor fraction estimation on postoperative plasma cfDNA WGS data (mean sequencing depth ~28x) from 90 stage III colorectal cancer patients with three years of follow-up. The tumor fraction (TF) estimates agreed with the available clinical information and ctDNA was detected in 77% (17/22) of recurring patients with a median lead time of 8 months compared to radiological imaging. We further validated ctDNAmers tumor fraction estimates based on a comparison with the mean cfDNA allele frequencies of somatic clonal SNVs identified from aligned primary tumor sequencing data. The TF estimates showed a strong Pearson correlation of 0.897 with the mean allele frequencies and improved ctDNA detection results across samples with an AUC of 0.79 compared to 0.75 if the mean allele frequency of clonal mutations is used.

bioinformatics↗

Timing and clustering co-occurring genome amplifications incancers

Clonal evolution in cancer is driven by genomic alterations that accumulate over time, shaping tumour progression, therapy resistance, and metastasis. Among these somatic events, genomic amplifications are a broad class of copy number alterations (CNAs) that can be mathematically timed (i.e., mapped to an abstract timeline). Existing methods successfully order amplifications in time but fail to understand their co-occurrence patterns. This limitation makes it harder to understand abrupt shifts of clonal and selection dynamics possibly linked to clones that acquire profound mutant genotypes and hold the potential to establish a novel evolutionary lineage. Here, we introduce TickTack, a hierarchical Bayesian mixture model for reconstructing the temporal order of copy number amplifications across the genome while simultaneously detecting co-occurrent events, offering a more comprehensive view of tumour evolutionary dynamics. This new model allows us to determine whether copy number amplifications accumulate gradually over multiple generations or occur in rapid succession within short time frames, providing deeper insights into genomic instability and tumor progression beyond traditional linear models. We validated our approach with synthetic data under various uncertainty settings and against competing approaches. Applying TickTack to 2,777 samples from the Pan-Cancer Analysis of Whole Genomes (PCAWG) project, a comprehensive resource spanning 38 tumor types, we inferred the temporal order of copy number amplifications, identifying cancer-specific co-occurring events. Our analysis revealed associations between early chromosomal instability and key driver mutations (TP53, BRCA1/2) in Esophageal Adenocarcinoma and uncovered recurrent evolutionary trajectories shaped by focal and arm-level copy gains. These findings highlight the role of saltational evolution in tumorigenesis and provide insights into genomic instability with possible implications for prognosis and targeted therapies. AvailabilitytickTack is available as an R package at https://caravagnalab.github.io/tickTack/ and the code to replicate the analysis is available from https://zenodo.org/records/14870458.

cancer biology↗

Model-based Bayesian inference of cancer dynamics from heterogenous longitudinal data

The kinetic parameters of cancer population dynamics are critical for developing reliable predictors of tumour growth patterns, extracting metrics for patient stratification and creating algorithms that can forecast clinically significant events. Here, we introduce a model-based Bayesian framework that leverages longitudinal phenotypic (e.g., tumour volume, cell counts) or genotypic (e.g., mutation frequency) data to infer critical parameters of tumour progression within a single patient. Our models uses population genetics to estimate probability distributions for tumour growth rates, initiation and extinction times, pinpointing abrupt shifts in tumour dynamics due to treatment response and revealing associations between drug resistance and pre-existing cancer cell populations. We apply our framework to address pivotal clinical questions across three major cancer types. In colorectal cancer, we use tumour markers data to identify extensive pre-existing RAS-linked resistance to cetuximab. In lung cancer, we use somatic mutation frequencies in citculating tumour DNA to determine prognostic growth rates and develop a test for monitoring minimal residual disease. In chronic lymphocytic leukaemia, we use white blood cell counts to stratify patients by growth patterns and predict time to treatment, advancing adaptive monitoring strategies.

cancer biology↗