Search bioRxivSearch

bioRxiv · 10.1101/000042

Routes for breaching and protecting genetic privacy

Abstract

We are entering the era of ubiquitous genetic information for research, clinical care, and personal curiosity. Sharing these datasets is vital for rapid progress in understanding the genetic basis of human diseases. However, one growing concern is the ability to protect the genetic privacy of the data originators. Here, we technically map threats to genetic privacy and discuss potential mitigation strategies for privacy-preserving dissemination of genetic data.\n\nAbout the AuthorsYaniv Erlich is a Fellow at the Whitehead Institute for Biomedical Research. Erlich received his Ph.D. from Cold Spring Harbor Laboratory in 2010 and B.Sc. from Tel-Aviv University in 2006. Prior to that, Erlich worked in computer security and was responsible for conducting penetration tests on financial institutes and commercial companies. Dr. Erlichs research involves developing new algorithms for computational human genetics.\n\nArvind Narayanan is an Assistant Professor in the Department of Computer Science and the Center for Information Technology and Policy at Princeton. He studies information privacy and security. His research has shown that data anonymization is broken in fundamental ways, for which he jointly received the 2008 Privacy Enhancing Technologies Award. His current research interests include building a platform for privacy-preserving data sharing.\n\nSummaryO_LIBroad data dissemination is essential for advancements in genetics, but also brings to light concerns regarding privacy.\nC_LIO_LIPrivacy breaching techniques work by cross-referencing two or more pieces of information to gain new, potentially undesirable knowledge on individuals or their families.\nC_LIO_LIBroadly speaking, the main routes to breach privacy are identity tracing, attribute disclosure, and completion of sensitive DNA information.\nC_LIO_LIIdentity tracing exploits quasi-identifiers in the DNA data or metadata to uncover the identity of an unknown genetic dataset.\nC_LIO_LIAttribute disclosure techniques work on known DNA datasets. They use the DNA information to link the identity of a person with a sensitive phenotype.\nC_LIO_LICompletion techniques also work on known DNA data. They try to uncover sensitive genomic areas that were masked to protect the participant.\nC_LIO_LIIn the last few years, we have witnessed a rapid growth in the range of techniques and tools to conduct these privacy-breaching attacks. Currently, most of the techniques are beyond the reach of the general public, but can be executed by trained persons with varying degrees of effort.\nC_LIO_LIThere is considerable debate regarding risk management. One camp supports a pragmatic, ad-hoc approach of privacy by obscurity and the other supports a systematic, mathematically-backed approach of privacy by design.\nC_LIO_LIPrivacy by design algorithms include access control, differential privacy, and cryptographic techniques. So far, data custodians of genetic databases mainly adopted access control as a mitigation strategy.\nC_LIO_LINew developments in cryptographic techniques may usher in an additional arsenal of security by design techniques.\nC_LI

Source connections

Explore related subjects

Keep this discovery

BibTeXRIS

Yaniv Erlich, Arvind Narayanan. 2013-11-07. Routes for breaching and protecting genetic privacy. https://doi.org/10.1101/000042

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Identification of genetic variants affecting vitamin D receptor binding and associations with autoimmune disease

Large numbers of statistically significant associations between sentinel SNPs and case-control status have been replicated by genome-wide association studies. Nevertheless, few underlying molecular mechanisms of complex disease are currently known. We investigated whether variation in binding of a transcription factor, the vitamin D receptor (VDR) whose activating ligand vitamin D has been proposed as a modifiable factor in multiple disorders, could explain any of these associations. VDR modifies gene expression by binding DNA as a heterodimer with the Retinoid X receptor (RXR).\n\nWe identified 43,332 genetic variants significantly associated with altered VDR binding affinity (VDR-BVs) using a high-resolution (ChIP-exo) genome-wide analysis of 27 HapMap lymphoblastoid cell lines. VDR-BVs are enriched in consensus RXR::VDR binding motifs, yet most fell outside of these motifs, implying that genetic variation often affects binding affinity only indirectly. Finally, we compared 341 VDR-BVs replicating by position in multiple individuals against background sets of variants lying within VDR-binding regions that had been matched in allele frequency and were independent with respect to linkage disequilibrium. In this stringent test, these replicated VDR-BVs were significantly (q < 0.1) and substantially (> 2-fold) enriched in genomic intervals associated with autoimmune and other diseases, including inflammatory bowel disease, Crohns disease and rheumatoid arthritis. The approachs validity is underscored by RXR::VDR motif sequence being predictive of binding strength and being evolutionarily constrained.\n\nOur findings are consistent with altered RXR::VDR binding contributing to immunity-related diseases. Replicated VDR-BVs associated with these disorders could represent causal disease risk alleles whose effect may be modifiable by vitamin D levels.

Genomics

Two novel genes discovered in human mitochondrial DNA using PacBio full-length transcriptome data

In this study, we introduced a general framework to use PacBio full-length transcriptome sequencing for the investigation of the fundamental problems in mitochondrial biology, e.g. genome arrangement, heteroplasmy, RNA processing and the regulation of transcription or replication. As a result, we produced the first full-length human mitochondrial transcriptome from the MCF7 cell line based on the PacBio platform and characterized the human mitochondrial transcriptome with more comprehensive and accurate information. The most important finding was two novel lnRNAs hsa-MDL1 and hsa-MDL1AS, which are encoded by the mitochondrial D-loop regions. We propose hsa-MDL1 and hsa-MDL1AS, as the precursors of transcription initiation RNAs (tiRNAs), belong to a novel class of long non-coding RNAs (lnRNAs), which is named as long tiRNAs (ltiRNAs). Based on the mitochondrial RNA processing model, the primary tiRNAs, precursors and mature tiRNAs could be discovered to completely reveal tiRNAs from their origins to functions. The MDL1 and MDL1AS lnRNAs and their regulation mechanisms exist ubiquitously from insects to human.

Genomics

Omics and bioinformatics approaches to target boar taint

In livestock species, a rapid growth in high-throughput omics data has accelerated the pace of studies that target to dissect economically important traits to provide better quality animal products to consumers. In pig industries, young boars are generally castrated to remove boar taint, a phenotypic and inheritable trait well-known by an abnormally bad smell and taste in pork meat derived from some uncastrated male pigs. Existence of porcine reference genome made possible to catalogue genome-wide QTLs, candidate genes and biomarkers in associations with boar taint and other industrially significant traits in pigs. The aim of this paper to review the contribution of bioinformatics resources and omics technology in boar taint related studies. This paper also provides concise details about state-of-the-art sequencing technology.

Genomics