Search bioRxiv⌕ Search

Biology subjects

Pentyala, S.

Publications and source records attributed to Pentyala, S..

2 recordsLinked to original sources

Towards Useful and Private Synthetic Omics: Community Benchmarking of Generative Models for Transcriptomics Data

BackgroundThe synthesis of anonymized data derived from real-world cohorts offers a promising strategy for regulatory-compliant and privacy-preserving biological data sharing, potentially facilitating model development that can improve predictive performance. However, the extent to which generative models can preserve biological signals while remaining resilient to adversarial privacy attacks in high-dimensional omics contexts remains underexplored. To address this gap, the CAMDA 2025 Health Privacy Challenge launched a community-driven effort to systematically benchmark synthetic and privacy-preserving data generation for bulk RNA-seq cohorts. ResultsBuilding on this initiative, we systematically benchmarked 11 generative methods across two cancer cohorts ([~]1,000 and [~]5,000 patients) over 978 landmark genes. Methods were evaluated across complementary axes of distributional fidelity, downstream utility, biological plausibility and empirical privacy risk, with emphasis on trade-offs between vulnerability to membership inference attacks (MIA) and other evaluation dimensions. Expressive deep generative models achieved strong predictive utility and differential expression recovery, but were often more vulnerable to membership inference risk. Differentially private methods improved resistance to attacks at the cost of reduced utility, while simpler statistical approaches offered competitive utility with moderate privacy risk and fast training. ConclusionsSynthetic bulk RNA-seq quality is inherently multi-dimensional and shaped by trade-offs between utility, biological preservation and privacy. Our results indicate that differences in model architecture drive distinct trade-offs across these axes, suggesting that model choice should align with dataset characteristics, intended downstream use and privacy requirements. Privacy risk should also be assessed using multiple complementary attack methods and, where possible, formal differential privacy protection.

bioinformatics↗

Privacy Vulnerabilities in Synthetic Single-Cell RNA-Sequence Data

Single-cell RNA sequencing (scRNA-seq) data is subject to strict access control due to its sensitive nature, motivating the use of synthetic data generation (SDG) for privacy-preserving data sharing. We present the first adversarial privacy attack that performs meaningfully above random guessing against state-of-the-art scRNA-seq SDG methods. Our attack enables donor-level membership inference, demonstrating that leading SDG techniques fail to adequately mask which individuals were used to train the generator. We show that privacy leakage increases as the number of training donors decreases. Although the attack is designed to exploit vulnerabilities in scDesign2, we find that it also succeeds against synthetic data generated by other leading methods, including scDesign3 and scVI. This transferability indicates that an adversary can infer sensitive information from synthetic data without access to the training procedure, model parameters, or even the underlying generation algorithm. Finally, we investigate the use of perturbation with noise during the SDG process as a first-line defense, empirically evaluating its effectiveness in neutralizing the attack and its impact on utility.

genomics↗