Search bioRxivSearch

Biology subjects

Murphy, S.

Publications and source records attributed to Murphy, S..

5 recordsLinked to original sources

Semi-supervised Encoding for Outlier Detection in Clinical Observation Data

Background and ObjectiveTo evaluate the utility of encoding for outlier detection in clinical observation data from Electronic Health Records (EHR).\n\nMethodsThis article presents a semi-supervise encoding approach (super-encoding) for constructing a non-linear exemplar data distribution from EHR data and detecting non-conforming observations as outliers. Two hypotheses are tested using experimental design and non-parametric hypothesis testing procedures to evaluate the outlier detection performance of the semi-supervised encoding approach and increasing demographic precision in encoding.\n\nResultsThe experiments involved applying 492 encoders to 30 laboratory tests extracted from the Research Patient Data Registry (RPDR) from Partners HealthCare. We report results obtained from 14,760 encoders. The semi-supervised encoders (super-encoders) outperformed conventional autoencoders in outlier detection. Adding age at observation to the baseline encoder (that only included observation value as the feature) slightly improved outlier detection. Top-nine performing encoders are introduced. The best outlier detection performance was from a semi-supervised encoder, with observation value as the single feature and a single hidden layer, built on one percent of the data and one percent reconstruction error. At least one encoder had a Youdens J index higher than 0.9999 for all 30 observations.\n\nConclusionGiven the multiplicity of distributions for a single observation in EHR data (i.e., same observation represented with different names or units), as well as non-linearity of human observations, encoding offers huge promises for outlier detection in large-scale data repositories.

bioinformatics

Spatial Capture-Recapture for Categorically Marked Populations with An Application to Genetic Capture-Recapture

Recently introduced unmarked spatial capture-recapture (SCR), spatial mark-resight (SMR), and 2-flank spatial partial identity models (SPIM) extend the domain of SCR to populations or observation systems that do not always allow for individual identity to be determined with certainty. For example, some species do not have natural marks that can reliably produce individual identities from photographs, and some methods of observation produce partial identity samples as is the case with remote cameras that sometimes produce single flank photographs. These models share the feature that they probabilistically resolve the uncertainty in individual identity using the spatial location where samples were collected. Spatial location is informative of individual identity in spatially structured populations with home range sizes smaller than the extent of the trapping array because a latent identity sample is more likely to have been produced by an individual living near the trap where it was recorded than an individual living further away from the trap. Further, the level of information about individual identity that a spatial location contains is determined by two key ecological concepts, population density and home range size. The number of individuals that could have produced a latent or partial identity sample increases as density and home range size increase because more individual home ranges will overlap any given trap. We show this uncertainty can be quantified using a metric describing the expected magnitude of uncertainty in individual identity for any given population density and home range size, the Identity Diversity Index (IDI). We then show that the performance of latent and partial identity SCR models varies as a function of this index and produces imprecise and biased estimates in many high IDI scenarios when data are sparse. We then extend the unmarked SCR model to incorporate partially identifying covariates which reduce the level of uncertainty in individual identity, increasing the reliability and precision of density estimates, and allowing reliable density estimation in scenarios with higher IDI values and with more sparse data. We illustrate the performance of this \"categorical SPIM\" via simulations and by applying it to a black bear data set using microsatellite loci as categorical covariates, where we reproduce the full data set estimates with only slightly less precision using fewer loci than necessary for confident individual identification. The categorical SPIM offers an alternative to using probability of identity criteria for classifying genotypes as unique, shifting the \"shadow effect\", where more than one individual in the population has the same genotype, from a source of bias to a source of uncertainty. We discuss the difficulties that real world data sets pose for latent identity SCR methods, most importantly, individual heterogeneity in detection function parameters, and argue that the addition of partial identity information reduces these concerns. We then discuss how the categorical SPIM can be applied to other wildlife sampling scenarios such as remote camera surveys, where natural or researcher-applied partial marks can be observed in photographs. Finally, we discuss how the categorical SPIM can be added to SMR, 2-flank SPIM, or other future latent identity SCR models.

ecology

Genetic validation of bipolar disorder identified by automated phenotyping using electronic health records

Bipolar disorder (BD) is a heritable mood disorder characterized by episodes of mania and depression. Although genomewide association studies (GWAS) have successfully identified genetic loci contributing to BD risk, sample size has become a rate-limiting obstacle to genetic discovery. Electronic health records (EHRs) represent a vast but relatively untapped resource for high-throughput phenotyping. As part of the International Cohort Collection for Bipolar Disorder (ICCBD), we previously validated automated EHR-based phenotyping algorithms for BD against in-person diagnostic interviews (Castro et al. 2015). Here, we establish the genetic validity of these phenotypes by determining their genetic correlation with traditionally-ascertained samples. Case and control algorithms were derived from structured and narrative text in the Partners Healthcare system comprising more than 4.6 million patients over 20 years. Genomewide genotype data for 3,330 BD cases and 3,952 controls of European ancestry were used to estimate SNP-based heritability (h2g) and genetic correlation(rg) between EHR-based phenotype definitions and traditionally-ascertained BD cases in GWAS by the ICCBD and Psychiatric Genomics Consortium (PGC) using LD score regression. We evaluated BD cases identified using 4 EHR-based algorithms: an NLP-based algorithm (95-NLP) and 3 rule-based algorithms using codified EHR with decreasing levels of stringency - \"coded-strict\", \"coded-broad\", and \"coded-broad based on a single clinical encounter\" (coded-broad-SV). The analytic sample comprised 862 95-NLP, 1,968 coded-strict, 2,581 coded-broad, 408 coded-broad-SV BD cases, and 3,952 controls. The estimated h2g were 0.24 (p=0.015), 0.09 (p=0.064), 0.13 (p=0.003), 0.00 (p=0.591) for 95-NLP, coded-strict, coded-broad and coded-broad-SV BD, respectively. The h2g for all EHR-based cases combined except coded-broad-SV (excluded due to 0 h2g) was 0.12 (p=0.004). These h2g were lower or similar to the h2g observed by the ICCBD+PGCBD (0.23, p=3.17E-80, total N=33,181). However, the rg between ICCBD+PGCBD and the EHR-based cases were high for 95-NLP (0.66, p=3.69x10-5), coded-strict (1.00, p=2.40x10-4), and coded-broad (0.74, p=8.11x10-7). The rg between EHR-based BDs ranged from 0.90 to 0.98. These results provide the first genetic validation of automated EHR-based phenotyping for BD and suggest that this approach identifies cases that are highly genetically correlated with those ascertained through conventional methods. High throughput phenotyping using the large data resources available in EHRs represents a viable method for accelerating psychiatric genetic research.

genetics

Cadmium Exposure Increases The Risk Of Juvenile Obesity: A Human And Zebrafish Comparative Study

OBJECTIVEHuman obesity is a complex metabolic disorder disproportionately affecting people of lower socioeconomic strata, and ethnic minorities, especially African Americans and Hispanics. Although genetic predisposition and a positive energy balance are implicated in obesity, these factors alone do not account for the excess prevalence of obesity in lower socioeconomic populations. Therefore, environmental factors, including exposure to pesticides, heavy metals, and other contaminants, are agents widely suspected to have obesogenic activity, and they also are spatially correlated with lower socioeconomic status. Our study investigates the causal relationship between exposure to the heavy metal, cadmium (Cd), and obesity in a cohort of children and a zebrafish model of adipogenesis.\n\nDESIGNAn extensive collection of first trimester maternal blood samples obtained as part of the Newborn Epigenetics Study (NEST) were analyzed for the presence Cd, and these results were cross analyzed with the weight-gain trajectory of the children through age five years. Next, the role of Cd as a potential obesogen was analyzed in an in vivo zebrafish model.\n\nRESULTSOur analysis indicates that the presence of Cd in maternal blood during pregnancy is associated with increased risk of juvenile obesity in the offspring, independent of other variables, including lead (Pb) and smoking status. Our results are recapitulated in a zebrafish model, in which exposure to Cd at levels approximating those observed in the NEST study is associated with increased adiposity.\n\nCONCLUSIONOur findings identify Cd as potential human obesogen. Moreover, these observations are recapitulated in a zebrafish model, suggesting that the underlying mechanisms may be evolutionarily conserved, and that zebrafish may be a valuable model for uncovering pathways leading to Cd-mediated obesity in human populations.

epidemiology

The RS Domain of Human CFIm68 Plays a Key Role in Selection Between Alternative Sites of Pre-mRNA Cleavage and Polyadenylation

Many eukaryotic protein-coding genes give rise to alternative mRNA isoforms with identical protein-coding capacities but which differ in the extents of their 3{acute} untranslated regions (3{acute}UTRs), due to the usage of alternative sites of pre-mRNA cleavage and polyadenylation. By governing the presence of regulatory 3{acute}UTR sequences, this type of alternative polyadenylation (APA) can significantly influence the stability, localisation and translation efficiency of mRNA. Though a variety of molecular mechanisms for APA have been proposed, previous studies have identified a pivotal role for the multi-subunit cleavage factor I (CFIm) in this process in mammals. Here we show that, in line with previous reports, depletion of the CFIm 68 kDa subunit (CFIm68) by CRISPR/Cas9-mediated gene disruption in HEK293 cells leads to a shift towards the use of promoter-proximal poly(A) sites. Using these cells as the basis for a complementation assay, we show that CFIm68 lacking its arginine/serine-rich (RS) domain retains the ability to form a nuclear complex with other CFIm subunits, but selectively lacks the capacity to restore polyadenylation at promoter-distal sites. In addition, nanoparticle-mediated analysis indicates that the RS domain is extensively phosphorylated in vivo. Overall, these results suggest that the CFIm68 RS domain makes a key regulatory contribution to APA.

molecular biology