Search bioRxivSearch

Biology subjects

Do, R.

Publications and source records attributed to Do, R..

5 recordsLinked to original sources

Unsupervised Machine learning to subtype Sepsis-Associated Acute Kidney Injury

ObjectiveAcute kidney injury (AKI) is highly prevalent in critically ill patients with sepsis. Sepsis-associated AKI is a heterogeneous clinical entity, and, like many complex syndromes, is composed of distinct subtypes. We aimed to agnostically identify AKI subphenotypes using machine learning techniques and routinely collected data in electronic health records (EHRs).\n\nDesignCohort study utilizing the MIMIC-III Database.\n\nSettingICUs from tertiary care hospital in the U.S.\n\nPatientsPatients older than 18 years with sepsis and who developed AKI within 48 hours of ICU admission.\n\nInterventionsUnsupervised machine learning utilizing all available vital signs and laboratory measurements.\n\nMeasurements and Main ResultsWe identified 1,865 patients with sepsis-associated AKI. Ten vital signs and 691 unique laboratory results were identified. After data processing and feature selection, 59 features, of which 28 were measures of intra-patient variability, remained for inclusion into an unsupervised machine-learning algorithm. We utilized k-means clustering with k ranging from 2 - 10; k=2 had the highest silhouette score (0.62). Cluster 1 had 1,358 patients while Cluster 2 had 507 patients. There were no significant differences between clusters on age, race or gender. We found significant differences in comorbidities and small but significant differences in several laboratory variables (hematocrit, bicarbonate, albumin) and vital signs (systolic blood pressure and heart rate). In-hospital mortality was higher in cluster 2 patients, 25% vs. 20%, p=0.008. Features with the largest differences between clusters included variability in basophil and eosinophil counts, alanine aminotransferase levels and creatine kinase values.\n\nConclusionsUtilizing routinely collected laboratory variables and vital signs in the EHR, we were able to identify two distinct subphenotypes of sepsis-associated AKI with different outcomes. Variability in laboratory variables, as opposed to their actual value, was more important for determination of subphenotypes. Our findings show the potential utility of unsupervised machine learning to better subtype AKI.

bioinformatics

The Genetic Landscape of Diamond-Blackfan Anemia

Diamond-Blackfan anemia (DBA) is a rare bone marrow failure disorder that affects 1 in 100,000 to 200,000 live births and has been associated with mutations in components of the ribosome. In order to characterize the genetic landscape of this genetically heterogeneous disorder, we recruited a cohort of 472 individuals with a clinical diagnosis of DBA and performed whole exome sequencing (WES). Overall, we identified rare and predicted damaging mutations in likely causal genes for 78% of individuals. The majority of mutations were singletons, absent from population databases, predicted to cause loss of function, and in one of 19 previously reported genes encoding for a diverse set of ribosomal proteins (RPs). Using WES exon coverage estimates, we were able to identify and validate 31 deletions in DBA associated genes. We also observed an enrichment for extended splice site mutations and validated the diverse effects of these mutations using RNA sequencing in patientderived cell lines. Leveraging the size of our cohort, we observed several robust genotype-phenotype associations with congenital abnormalities and treatment outcomes. In addition to comprehensively identifying mutations in known genes, we further identified rare mutations in 7 previously unreported RP genes that may cause DBA. We also identified several distinct disorders that appear to phenocopy DBA, including 9 individuals with biallelic CECR1 mutations that result in deficiency of ADA2. However, no new genes were identified at exome-wide significance, suggesting that there are no unidentified genes containing mutations readily identified by WES that explain > 5% of DBA cases. Overall, this comprehensive report should not only inform clinical practice for DBA patients, but also the design and analysis of future rare variant studies for heterogeneous Mendelian disorders.

genetics

The landscape of pervasive horizontal pleiotropy in human genetic variation is driven by extreme polygenicity of human traits and diseases

Horizontal pleiotropy, where one variant has independent effects on multiple traits, is important for our understanding of the genetic architecture of human phenotypes. We develop a method to quantify horizontal pleiotropy using genome-wide association summary statistics and apply it to 372 heritable phenotypes measured in 361,194 UK Biobank individuals. Horizontal pleiotropy is pervasive throughout the human genome, prominent among highly polygenic phenotypes, and enriched in active regulatory regions. Our results highlight the central role horizontal pleiotropy plays in the genetic architecture of human phenotypes. The HOrizontal Pleiotropy Score (HOPS) method is available on Github at https://github.com/rondolab/HOPS.

genetics

Genetic Diversity Turns a New PAGE in Our Understanding of Complex Traits

Summary/AbstractGenome-wide association studies (GWAS) have laid the foundation for investigations into the biology of complex traits, drug development, and clinical guidelines. However, the dominance of European-ancestry populations in GWAS creates a biased view of the role of human variation in disease, and hinders the equitable translation of genetic associations into clinical and public health applications. The Population Architecture using Genomics and Epidemiology (PAGE) study conducted a GWAS of 26 clinical and behavioral phenotypes in 49,839 non-European individuals. Using strategies designed for analysis of multi-ethnic and admixed populations, we confirm 574 GWAS catalog variants across these traits, and find 38 secondary signals in known loci and 27 novel loci. Our data shows strong evidence of effect-size heterogeneity across ancestries for published GWAS associations, substantial benefits for fine-mapping using diverse cohorts, and insights into clinical implications. We strongly advocate for continued, large genome-wide efforts in diverse populations to reduce health disparities.

genetics

Widespread pleiotropy confounds causal relationships between complex traits and diseases inferred from Mendelian randomization

A fundamental assumption in inferring causality of an exposure on complex disease using Mendelian randomization (MR) is that the genetic variant used as the instrumental variable cannot have pleiotropic effects. Violation of this no pleiotropy assumption can cause severe bias. Emerging evidence have supported a role for pleiotropy amongst disease-associated loci identified from GWA studies. However, the impact and extent of pleiotropy on MR is poorly understood. Here, we introduce a method called the Mendelian Randomization Pleiotropy RESidual Sum and Outlier (MR-PRESSO) test to detect and correct for pleiotropy in multi-instrument summary-level MR testing. We show using simulations that existing approaches are less sensitive to the detection of pleiotropy when it occurs in a subset of instrumental variables, as compared to MR-PRESSO. Next, we show that pleiotropy is widespread in MR, occurring in 41% amongst significant causal relationships (out of 4,250 MR tests total) from pairwise comparisons of 82 complex traits and diseases from summary level genome-wide association data. We demonstrate that pleiotropy causes distortion between-168% and 189% of the causal estimate in MR. Furthermore, pleiotropy induces false positive causal relationships-defined as those causal estimates that were no longer statistically significant in the pleiotropy corrected MR test but were previously significant in the naive MR test-in up to 10% of the MR tests using a P < 0.05 cutoff that is commonly used in MR studies. Finally, we show that MR-PRESSO can correct for distortion in the causal estimate in most cases. Our results demonstrate that pleiotropy is widespread and pervasive, and must be properly corrected for in order to maintain the validity of MR.

genomics