Search bioRxiv⌕ Search

bioRxiv · 10.64898/2026.06.10.731306

Evaluating anonymized genome re-identification using polygenic predictions and its implications for data privacy

Abstract

Re-identification by phenotypic prediction aims to determine whether a genome belongs to a specific individual by comparing the individuals known traits with those predicted from the genome. This type of tracing attack is widely discussed in the genomic privacy literature, yet previous studies have been criticized for overstating its practical risks. Over the past decade, genome-wide association studies (GWAS) with increasing sample size improved the accuracy of phenotypic prediction, potentially enhancing such attacks. To quantify their real-world threat, we developed a probabilistic framework that estimates the likelihood of a match between an individuals observed traits and polygenic scores (PGS) derived from a genome, while accounting for prediction accuracy and genetic and environmental correlations between the traits. We benchmarked this re-identification method and examined how the prior probability (reflecting the a priori chance that a random genome and set of traits correspond to the same person) affects performance. Finally, we assessed whether sensitive information could be inferred through this attack by attempting to predict multiple sensitive haplotypes, such as APOE-{varepsilon}4 (linked with Alzheimers disease). Our re-identification method outperformed a state-of-the-art tool, and reached a precision above 99% for a recall of 40% when considering a prior of 50%. However, after considering real-world settings, we estimated that realistic priors would not exceed 4 x 10-4%, resulting in a precision lower than 0.13% at the same recall (40%). The inference of sensitive genotypes also proved ineffective, as achieving a precision above 50% for identifying APOE-{varepsilon}4 carriers was only possible at a recall below 20%. To conclude, although re-identification by phenotypic prediction is technically feasible, our findings indicate that its effectiveness in real-world conditions is limited. These results counterpoint to earlier claims of severe genomic privacy risks and offer guidance for policymakers, biobank administrators, and research participants.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Cavinato, T., Hofmeister, R. J., Kutalik, Z.. 2026-06-10. Evaluating anonymized genome re-identification using polygenic predictions and its implications for data privacy. https://doi.org/10.64898/2026.06.10.731306

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Synergistic Variants in C-terminal Binding Protein 1 and Alkaline Phosphatase Lead to Mandibular Hypoplasia Through Impaired Wnt Signaling: An Oligogenic Model

Craniofacial malformations account for one third of all congenital anomalies. Genetic factors play a vital role, yet the list of causal genes and their mechanisms are far from complete. As part of a larger effort to sequence patients with micrognathia and Pierre-Robin sequence, we identified two candidate pathogenic missense variants in C-terminal binding protein 1 (CTBP1) along with a heterozygous early stop missense variant in alkaline phosphatase (ALPL) in a proband with mandibular hypoplasia. Ctbp1 has been shown to regulate Wnt/{beta}-Catenin signaling but it has not yet been implicated in craniofacial development. Here we generated two orthologous variants of Ctbp1 mimicking the patient variants using genome editing in mice and explored the micrognathia phenotype in combination with a previously reported Alpl null allele. Ctbp1Q148H/G238S; Alplnull/Wt complex heterozygous mutants have smaller mandibles recapitulating the human mandibular hypoplasia. We identified that a reduction in cell proliferation and active {beta}-Catenin levels could possibly account for the micrognathia phenotype in the Ctbp1; Alpl complex heterozygous. These data uncover a novel role for Ctbp1 in craniofacial development and highlight the complex genetic and molecular signaling in the pathogenesis of craniofacial malformations.

genetics↗

Intergenerational instability of the C9orf72 hexanucleotide repeat

The C9orf72 hexanucleotide repeat expansion (HRE) is the most common cause of amyotrophic lateral sclerosis (ALS) and frontotemporal dementia (FTD). It follows autosomal dominant inheritance in families, however, a high proportion of cases are sporadic, raising the possibility of parental premutation. We have demonstrated that intermediate-length alleles (IAs) with >18 repeats (allele frequency ~1%) belong to the same pool of haplotypes as the HRE, suggesting shared ancestry. Here, we tested whether alleles with >18 repeats expand in parental transmission. We used two repeat-primed PCR methods to analyze allele lengths in 539 genetically unselected parent-offspring pairs and in 152 pairs known to carry the SNP (rs139185008*C) that tags >18 repeat IAs and the HRE in Finland. We discovered intergenerational repeat length changes only in >20 repeat alleles. A significant (P = 0.0059) sex bias in 6-40 repeat alleles was noted using a logistic regression model. In this allele range, 12 out of 16 expansions were paternally inherited and 6 out of 7 contractions were maternally inherited. The expansion rate of 20-40 repeat alleles was 34 % in paternal and 11 % in maternal transmissions. In the 20-40 repeat range, most intergenerational expansions were 1-4 repeats in size (15/16), but one larger jump, a paternal expansion from 27 to 73 repeats, was observed. These results demonstrate that alleles with >20 repeats have an increased likelihood of instability, that a paternal expansion bias is observed in alleles with 20-40 repeats, and that expansion events are predominantly 1-4 repeats in size.

genetics↗

Unravelling the role of IRX4 variants in non-syndromic and Down syndrome associated congenital heart disease

IRX4 is a TALE- homeodomain transcription factor which is essential for cardiac development. In murine models, Irx4 deficiency leads to impaired ventricular function and results in cardiomyopathy. To elucidate the role of IRX4 in human congenital heart disease (CHD), Sanger sequencing of the IRX4 gene was performed in 205 individuals with non-syndromic CHD, 24 Down syndrome (DS) cases with CHD, 27 DS cases without CHD, and 150 healthy control individuals. Two novel (p.Ser24Asn and p.Thr217Iso) and one reported variant (rs2232376) were identified in non-syndromic CHD. Concurrently, rs2232376 was also detected in DS with CHD. The first novel (p.Ser24Asn) and reported (rs2232376) variants lie in the N-terminal region while the second novel (Thr217Iso) variant lies within the TALE homeodomain. In silico structural modelling suggested that both the novel variants (p.Ser24Asn and Thr217Iso) induce conformational changes in the IRX4 protein, potentially altering its DNA-binding affinity. A significant reduced expression of IRX4 muteins was noted in Western blotting by both variants (p.Ser24Asn and Thr217Iso). Furthermore, luciferase reporter assays demonstrated decline in the activity of Nanog promoter and HEY2 enhancer in response to both the variants which was further corroborated by decrease mRNA expression in qRT-PCR. Additional downstream targets, including Nfyc, Nppa, and Bmp10, also exhibited anomalous expression due to both the variants (p.Ser24Asn and Thr217Iso). Altogether, the aberrant expression of muteins as well as downstream target genes along with compromised activities of promoters substantiate the pathogenic potential of the identified IRX4 variants and underscore the critical role of IRX4 in regulating multiple stages of cardiogenesis.

genetics↗