Search bioRxivSearch

Biology subjects

McKinney, B. A.

Publications and source records attributed to McKinney, B. A..

2 recordsLinked to original sources

A nonlinear simulation framework supports adjusting for age when analyzing BrainAGE.

Several imaging modalities, including T1-weighted structural imaging, diffusion tensor imaging, and functional MRI can show chronological age related changes. Employing machine learning algorithms, an individuals imaging data can predict their age with reasonable accuracy. While details vary according to modality, the general strategy is to: 1) extract image-related features, 2) build a model on a training set that uses those features to predict an individuals age, 3) validate the model on a test dataset, producing a predicted age for each individual, 4) define the \"Brain Age Gap Estimate\" (BrainAGE) as the difference between an individuals predicted age and his/her chronological age, and 5) estimate the relationship between BrainAGE and other variables of interest, and 6) make inferences about those variables and accelerated or delayed brain aging. For example, a group of individuals with overall positive BrainAGE may show signs of accelerated aging in other variables as well. There is inevitably an overestimation of the age of younger individuals and an underestimation of the age of older individuals due to regression to the mean. The correlation between chronological age and BrainAGE may significantly impact the relationship between BrainAGE and other variables of interest when they are also related to age. In this study, we examine the detectability of variable effects under different assumptions. We use empirical results from two separate datasets [training=475 healthy volunteers, aged 18 - 60 years (259 female); testing=489 participants including people with mood/anxiety, substance use, eating disorders and healthy controls, aged 18 - 56 years (312 female)] to inform simulation parameter selection. Outcomes in simulated and empirical data strongly support the proposal that models incorporating BrainAGE should include chronological age as a covariate. We propose either including age as a covariate in step 5 of the above framework, or employing a multistep procedure where age is regressed on BrainAGE prior to step 5, producing BrainAGE Residualized (BrainAGER) scores.

bioinformatics

Statistical Inference Relief (STIR) feature selection

MotivationRelief is a family of machine learning algorithms that uses nearest-neighbors to select features whose association with an outcome may be due to epistasis or statistical interactions with other features in high-dimensional data. Relief-based estimators are non-parametric in the statistical sense that they do not have a parameterized model with an underlying probability distribution for the estimator, making it difficult to determine the statistical significance of Relief-based attribute estimates. Thus, a statistical inferential formalism is needed to avoid imposing arbitrary thresholds to select the most important features.\n\nMethodsWe reconceptualize the Relief-based feature selection algorithm to create a new family of STatistical Inference Relief (STIR) estimators that retains the ability to identify interactions while incorporating sample variance of the nearest neighbor distances into the attribute importance estimation. This variance permits the calculation of statistical significance of features and adjustment for multiple testing of Relief-based scores. Specifically, we develop a pseudo t-test version of Relief-based algorithms for case-control data.\n\nResultsWe demonstrate the statistical power and control of type I error of the STIR family of feature selection methods on a panel of simulated data that exhibits properties reflected in real gene expression data, including main effects and network interaction effects. We compare the performance of STIR when the adaptive radius method is used as the nearest neighbor constructor with STIR when thefixed-k nearest neighbor constructor is used. We apply STIR to real RNA-Seq data from a study of major depressive disorder and discuss STIRs straightforward extension to genome-wide association studies.\n\nAvailabilityCode and data available at http://insilico.utulsa.edu/software/STIR.\n\nContactbrett.mckinney@gmail.com

bioinformatics