Search bioRxivSearch

Biology subjects

Edge, M. D.

Publications and source records attributed to Edge, M. D..

3 recordsLinked to original sources

Attacks on genetic privacy via uploads to genealogical databases

Direct-to-consumer (DTC) genetics services are increasingly popular for genetic genealogy, with tens of millions of customers as of 2019. Several DTC genealogy services allow users to upload their own genetic datasets in order to search for genetic relatives. A user and a target person in the database are identified as genetic relatives if the users uploaded genome shares one or more sufficiently long segments in common with that of the target person--that is, if the two genomes share one or more long regions identical by state (IBS). IBS matches reveal some information about the genotypes of the target person, particularly if the chromosomal locations of IBS matches are shared with the uploader. Here, we describe several methods by which an adversary who wants to learn the genotypes of people in the database can do so by uploading multiple datasets. Depending on the methods used for IBS matching and the information about IBS segments returned to the user, substantial information about users genotypes can be revealed with a few hundred uploaded datasets. For example, using a method we call IBS tiling, we estimate that an adversary who uploads approximately 900 publicly available genomes could recover at least one allele at SNP sites across up to 82% of the genome of a median person of European ancestries. In databases that detect IBS segments using unphased genotypes, approximately 100 uploads of falsified datasets can reveal enough genetic information to allow accurate genome-wide imputation of every person in the database. Different DTC services use different methods for identifying and reporting IBS segments, leading to differences in vulnerability to the attacks we describe. We provide a proof-of-concept demonstration that the GEDmatch database in particular uses unphased genotypes to detect IBS and is vulnerable to genotypes being revealed by artificial datasets. We suggest simple-to-implement suggestions that will prevent the exploits we describe and discuss our results in light of recent trends in genetic privacy, including the recent use of uploads to DTC genetic genealogy services by law enforcement.

genomics

Assortative mating and the dynamical decoupling of genetic admixture levels from phenotypes that differ between source populations

Source populations for an admixed population can possess distinct patterns of genotype and pheno-type at the beginning of the admixture process. Such differences are sometimes taken to serve as markers of ancestry--that is, phenotypes that are initially associated with the ancestral background in one source population are taken to reflect ancestry in that population. Examples exist, however, in which genotypes or phenotypes initially associated with ancestry in one source population have decoupled from overall admixture levels, so that they no longer serve as proxies for genetic ancestry. We develop a mechanistic model for describing the joint dynamics of admixture levels and phenotype distributions in an admixed population. The approach includes a quantitative-genetic model that relates a phenotype to underlying loci that affect its trait value. We consider three forms of mating. First, individuals might assort in a manner that is independent of the overall genetic admixture level. Second, individuals might assort by a quantitative phenotype that is initially correlated with the genetic admixture level. Third, individuals might assort by the genetic admixture level itself. Under the model, we explore the relationship between genetic admixture level and phenotype over time, studying the effect on this relationship of the genetic architecture of the phenotype. We find that the decoupling of genetic ancestry and phenotype can occur surprisingly quickly, especially if the phenotype is driven by a small number of loci. We also find that positive assortative mating attenuates the process of dissociation in relation to a scenario in which mating is random with respect to genetic admixture and with respect to phenotype. The mechanistic framework suggests that in an admixed population, a trait that initially differed between source populations might be a reliable proxy for ancestry for only a short time, especially if the trait is determined by relatively few loci. The results are potentially relevant in admixed human populations, in which phenotypes that have a perceived correlation with ancestry might have social significance as ancestry markers, despite declining correlations with ancestry over time.\n\nAuthor SummaryAdmixed populations are populations that descend from two or more populations that had been separated for a long time at the beginning of the admixture process. The source populations typically possess distinct patterns of genotype and phenotype. Hence, early in the admixture process, phenotypes of admixed individuals can provide information about the extent to which these individuals possess ancestry in a specific source population. To study correlations between admixture levels and phenotypes that differ between source populations, we construct a genetic and phenotypic model of the dynamical process of admixture. Under the model, we show that correlations between admixture levels and these phenotypes dissipate over time--especially if the genetic architecture of the phenotypes involves only a small number of loci, or if mating in the admixed population is random with respect to both the admixture levels and the phenotypes. The result has the implication that a trait that once reflected ancestry in a specific source population might lose this ancestry correlation. As a consequence, in human populations, after a sufficient length of time, salient phenotypes that can have social meaning as ancestry markers might no longer bear any relationship to genome-wide genetic ancestry.

evolutionary biology

Linkage disequilibrium connects genetic records of relatives typed with disjoint genomic marker sets

In familial searching in forensic genetics, a query DNA profile is tested against a database to determine whether it represents a relative of a database entrant. We examine the potential for using linkage disequilibrium to identify pairs of profiles as belonging to relatives when the query and database rely on nonoverlapping genetic markers. Considering data on individuals genotyped with both microsatellites used in forensic applications and genome-wide SNPs, we find that ~30-32% of parent-offspring pairs and ~35-36% of sib pairs can be identified from the SNPs of one member of the pair and the microsatellites of the other. The method suggests the possibility of performing familial searches of microsatellite databases using query SNP profiles, or vice versa. It also reveals that privacy concerns arising from computations across multiple databases that share no genetic markers in common entail risks not only for database entrants, but for their close relatives as well.

genetics