PRISM-G: an interpretable privacy scoring method for assessing risk in synthetic human genome data
Synthetic genomic data promises broader access, but unresolved privacy risks persist. In Europe, these risks increasingly hinder cross-border use of national genomic resources due to limitations in trust and legal interoperability rather than scientific demand. At the same time, privacy risk is not uniformly distributed: leakage driven by relatedness structure and rare-variant uniqueness can disproportionately affect underrepresented or vulnerable populations when shared data are misused or linked with external resources, making transparent, domain-aware measurement of privacy exposure central to responsible governance. We introduce PRISM-G, a model-agnostic framework that quantifies privacy exposure in synthetic genome data cross three complementary components: proximity to real genomes in genetic-coordinate space, replay of familial or population-structure patterns, and trait-linked exposure via rare variants and membership-inference signals. These components are normalized and combined through a risk-averse aggregation into a single 0-100 PRISM-G score. We evaluated PRISM-G on synthetic cohorts generated by a generative adversarial network (GAN), a restricted Boltzmann machine (RBM), and a logic-based SAT-solver (Genomator). Our results show that privacy vulnerabilities concentrate along different axes across models and marker densities, demonstrating that a single similarity-based metric is sufficient to characterize genomic privacy risk.