Search bioRxivSearch

Biology subjects

Emani, P. S.

Publications and source records attributed to Emani, P. S..

2 recordsLinked to original sources

PLIGHT: A tool to assess privacy risk by inferring identifying characteristics from sparse, noisy genotypes

Single nucleotide polymorphisms (SNPs) from omics data carry a high risk of reidentification for individuals and their relatives. While the ability of thousands of SNPs (especially rare ones) to identify individuals has been repeatedly demonstrated, the ready availability of small sets of noisy genotypes - such as from environmental DNA samples or functional genomics data - motivated us to quantify their informativeness. Here, we present a computational tool suite, PLIGHT ("Privacy Leakage by Inference across Genotypic HMM Trajectories"), that employs population-genetics-based Hidden Markov Models of recombination and mutation to find piecewise alignment of small, noisy query SNP sets to a reference haplotype database. We explore cases where query individuals are either known to be in a database, or not, and consider a variety of queries, including simulated genotype "mosaics" (composites from 2 source individuals) and genotypes from swabs of coffee cups from a known individual. Using PLIGHT on a database with ~5,000 haplotypes, we find for common, noise-free SNPs that only ten are sufficient to identify individuals, ~20 can identify both components in two-individual simulated mosaics, and 20-30 can identify first-order relatives (parents, children, and siblings). Using noisy coffee-cup-derived SNPs, PLIGHT identifies an individual (within the database) using ~30 SNPs. Moreover, even when the individual is not in the database, local genotype matches allow for some phenotypic information leakage based on coarse-grained GWAS SNP imputation and polygenic risk scores. Overall, PLIGHT maximizes the identifying information content of sparse SNP sets through exact or partial matches to databases. Finally, by quantifying such privacy attacks, PLIGHT helps determine the value of selectively sanitizing released SNPs without explicit assumptions about underlying population membership or allele frequencies. To make this practical, we provide a sanitization tool to remove the most identifying SNPs from a query set.

bioinformatics

Broad transcriptomic dysregulation across the cerebral cortex in ASD

Classically, psychiatric disorders have been considered to lack defining pathology, but recent work has demonstrated consistent disruption at the molecular level, characterized by transcriptomic and epigenetic alterations.1-3 In ASD, upregulation of microglial, astrocyte, and immune signaling genes, downregulation of specific synaptic genes, and attenuation of regional gene expression differences are observed.1,2,4-6 However, whether these changes are limited to the cortical association areas profiled is unknown. Here, we perform RNA-sequencing (RNA-seq) on 725 brain samples spanning 11 distinct cortical areas in 112 ASD cases and neurotypical controls. We identify substantially more genes and isoforms that differentiate ASD from controls than previously observed. These alterations are pervasive and cortex-wide, but vary in magnitude across regions, roughly showing an anterior to posterior gradient, with the strongest signal in visual cortex, followed by parietal cortex and the temporal lobe. We find a notable enrichment of ASD genetic risk variants among cortex-wide downregulated synaptic plasticity genes and upregulated protein folding gene isoforms. Finally, using snRNA-seq, we determine that regional variation in the magnitude of transcriptomic dysregulation reflects changes in cellular proportion and cell-type-specific gene expression, particularly impacting L3/4 excitatory neurons. These results highlight widespread, genetically-driven neuronal dysfunction as a major component of ASD pathology in the cerebral cortex, extending beyond association cortices to involve primary sensory regions.

neuroscience