Search bioRxiv⌕ Search

Biology subjects

Zeeshan, S.

Publications and source records attributed to Zeeshan, S..

2 recordsLinked to original sources

Discovering biomarkers associated and predicting cardiovascular disease with high accuracy using a novel nexus of machine learning techniques for precision medicine

Personalized interventions are deemed vital given the intricate characteristics, advancement, inherent genetic composition, and diversity of cardiovascular diseases (CVDs). The appropriate utilization of artificial intelligence (AI) and machine learning (ML) methodologies can yield novel understandings of CVDs, enabling improved personalized treatments through predictive analysis and deep phenotyping. In this study, we proposed and employed a novel approach combining traditional statistics and a nexus of cutting-edge AI/ML techniques to identify significant biomarkers for our predictive engine by analyzing the complete transcriptome of CVD patients. After robust gene expression data pre-processing, we utilized three statistical tests (Pearson correlation, Chi-square test, and ANOVA) to assess the differences in transcriptomic expression and clinical characteristics between healthy individuals and CVD patients. Next, the Recursive Feature Elimination (RFE) classifier assigned rankings to transcriptomic features based on their relation to the case-control variable. The top ten percent of commonly observed significant biomarkers were evaluated using four unique ML classifiers (Random Forest, Support Vector Machine, Xtreme Gradient Boosting Decision Trees, and k-Nearest Neighbors). After optimizing hyperparameters, the ensembled models, which were implemented using a soft voting classifier, accurately differentiated between patients and healthy individuals. We have uncovered 18 transcriptomic biomarkers that are highly significant in the CVD population that were used to predict disease with up to 96% accuracy. Additionally, we cross-validated our results with clinical records collected from patients in our cohort. The identified biomarkers served as potential indicators for early detection of CVDs. With its successful implementation, our newly developed predictive engine provides a valuable framework for identifying patients with CVDs based on their biomarker profiles.

genomics↗

Investigating variant and expression of CVD genes associated phenotypes among high-risk Heart Failure patients

Cardiovascular disease (CVD) is a leading cause of premature mortality in the US and the world. CVD comprises of several complex and mostly heritable conditions, which range from myocardial infarction to congenital heart disease. Here, we report our findings from an integrative analysis of gene expression, disease-causing gene variants, and associated phenotypes among CVD populations, with a focus on high-risk Heart Failure (HF) patients. We built a cohort using electronic health records (EHR) of consented patients with available samples, and then performed high-throughput whole-genome and RNA sequencing (RNA-seq) of key genes responsible for HF and other CVD pathologies. We also incorporated a translational aspect to our study by integrating genomics findings with patient medical records. This involved linking ICD-10 codes with our gene expression and variant data to identify associations with HF and other CVDs. Our in-depth gene expression analysis revealed differentially expressed genes associated with HF (41 genes) and other CVDs (23 genes). Furthermore, a variant analysis of whole-genome sequence data of CVD patients identified genes with altered gene expression (FLNA, CST3, LGALS3, and HBA1) with functional and nonfunctional mutations in these genes. Our study highlights the importance of an integrative approach that leverages gene expression, genetic mutations, and clinical data that will allow the prioritization of key driver genes for complex diseases to improve personalized healthcare.

genomics↗