Search bioRxivSearch

Biology subjects

Abdi, H.

Publications and source records attributed to Abdi, H..

3 recordsLinked to original sources

Generalization of the minimum covariance determinant algorithm for categorical and mixed data types

The minimum covariance determinant (MCD) algorithm is one of the most common techniques to detect anomalous or outlying observations. The MCD algorithm depends on two features of multivariate data: the determinant of a matrix (i.e., geometric mean of the eigenvalues) and Mahalanobis distances (MD). While the MCD algorithm is commonly used, and has many extensions, the MCD is limited to analyses of quantitative data and more specifically data assumed to be continuous. One reason why the MCD does not extend to other data types such as categorical or ordinal data is because there is not a well-defined MD for data types other than continuous data. To address the lack of MCD-like techniques for categorical or mixed data we present a generalization of the MCD. To do so, we rely on a multivariate technique called correspondence analysis (CA). Through CA we can define MD via singular vectors and also compute the determinant from CAs eigenvalues. Here we define and illustrate a generalized MCD on categorical data and then show how our generalized MCD extends beyond categorical data to accommodate mixed data types (e.g., categorical, ordinal, and continuous). We illustrate this generalized MCD on data from two large scale projects: the Ontario Neurodegenerative Disease Research Initiative (ONDRI) and the Alzheimers Disease Neuroimaging Initiative (ADNI), with genetics (categorical), clinical instruments and surveys (categorical or ordinal), and neuroimaging (continuous) data. We also make R code and toy data available in order to illustrate our generalized MCD.

bioinformatics

Multivariate genotypic analyses that identify specific genotypes to characterize disease and control groups in ADNI

INTRODUCTIONGenetic contributions to Alzheimers Disease (AD) are likely polygenic and not necessarily explained by uniformly applied linear and additive effects. In order to better understand the genetics of AD, we require statistical techniques to address both polygenic and possible non-additive effects.\n\nMETHODSWe used partial least squares-correspondence analysis (PLS-CA)--a method designed to detect multivariate genotypic effects. We used ADNI-1 (N = 756) as a discovery sample with two forms of PLS-CA: diagnosis-based and ApoE-based. We used ADNI-2 (N= 791) as a validation sample with a diagnosis-based PLS-CA.\n\nRESULTSWith PLS-CA we identified some expected genotypic effects (e.g., APOE/TOMM40, and APP) and a number of new effects that include, for examples, risk-associated genotypes in RBFOX1 and GPC6 and control-associated genotypes in PTPN14 and CPNE5.\n\nDISCUSSIONThrough the use of PLS-CA, we were able to detect complex (multivariate, genotypic) genetic contributions to AD, which included many non-additive and non-linear risk and possibly protective effects.

genomics

The Latent Semantic Space and Corresponding Brain Regions of the Functional Neuroimaging Literature via NeuroSynth

The functional neuroimaging literature has become increasingly complex and thus difficult to navigate. This complexity arises from the rate at which new studies are published and from the terminology that varies widely from study-to-study and even more so from discipline-to-discipline. One way to investigate and manage this problem is to build a \"semantic space\" that maps the different vocabulary used in functional neuroimaging literature. Such a semantic space will also help identify the primary research domains of neuroimaging and their most commonly reported brain regions. In this work, we analyzed the multivariate semantic structure of abstracts in Neurosynth and found that there are six primary domains of the functional neuroimaging literature each with their own preferred reported brain regions. Our analyses also highlight possible semantic sources of reported brain regions within and across domains because some research topics (e.g., memory disorders, substance use disorder) use heterogeneous terminology. Furthermore, we highlight the growth and decline of the primary domains over time. Finally, we note that our techniques and results form the basis of a \"recommendation engine\" that could help readers better navigate the neuroimaging literature.

neuroscience