Search bioRxivSearch

Biology subjects

Van Vleck, T.

Publications and source records attributed to Van Vleck, T..

2 recordsLinked to original sources

Unsupervised Machine learning to subtype Sepsis-Associated Acute Kidney Injury

ObjectiveAcute kidney injury (AKI) is highly prevalent in critically ill patients with sepsis. Sepsis-associated AKI is a heterogeneous clinical entity, and, like many complex syndromes, is composed of distinct subtypes. We aimed to agnostically identify AKI subphenotypes using machine learning techniques and routinely collected data in electronic health records (EHRs).\n\nDesignCohort study utilizing the MIMIC-III Database.\n\nSettingICUs from tertiary care hospital in the U.S.\n\nPatientsPatients older than 18 years with sepsis and who developed AKI within 48 hours of ICU admission.\n\nInterventionsUnsupervised machine learning utilizing all available vital signs and laboratory measurements.\n\nMeasurements and Main ResultsWe identified 1,865 patients with sepsis-associated AKI. Ten vital signs and 691 unique laboratory results were identified. After data processing and feature selection, 59 features, of which 28 were measures of intra-patient variability, remained for inclusion into an unsupervised machine-learning algorithm. We utilized k-means clustering with k ranging from 2 - 10; k=2 had the highest silhouette score (0.62). Cluster 1 had 1,358 patients while Cluster 2 had 507 patients. There were no significant differences between clusters on age, race or gender. We found significant differences in comorbidities and small but significant differences in several laboratory variables (hematocrit, bicarbonate, albumin) and vital signs (systolic blood pressure and heart rate). In-hospital mortality was higher in cluster 2 patients, 25% vs. 20%, p=0.008. Features with the largest differences between clusters included variability in basophil and eosinophil counts, alanine aminotransferase levels and creatine kinase values.\n\nConclusionsUtilizing routinely collected laboratory variables and vital signs in the EHR, we were able to identify two distinct subphenotypes of sepsis-associated AKI with different outcomes. Variability in laboratory variables, as opposed to their actual value, was more important for determination of subphenotypes. Our findings show the potential utility of unsupervised machine learning to better subtype AKI.

bioinformatics

Genetic Identification Of A Common Collagen Disease In Puerto Ricans Via Identity-By-Descent Mapping In A Health System

Achieving confidence in the causality of a disease locus is a complex task that often requires supporting data from both statistical genetics and clinical genomics. Here we describe a combined approach to identify and characterize a genetic disorder that leverages distantly related patients in a health system and population-scale mapping. We utilize genomic data to uncover components of distant pedigrees, in the absence of recorded pedigree information, in the multi-ethnic BioMe biobank in New York City. By linking to medical records, we discover a locus associated with genetic relatedness that also underlies extreme short stature. We link the gene, COL27A1, with a little-known genetic disease, previously thought to be rare and recessive. We demonstrate that disease manifests in both heterozygotes and homozygotes, indicating a common collagen disorder impacting up to 2% of individuals of Puerto Rican ancestry, leading to a better understanding of the continuum of complex and Mendelian disease.

genomics