Search bioRxiv⌕ Search

Biology subjects

Abdelhalim, H.

Publications and source records attributed to Abdelhalim, H..

2 recordsLinked to original sources

Discovering biomarkers associated and predicting cardiovascular disease with high accuracy using a novel nexus of machine learning techniques for precision medicine

Personalized interventions are deemed vital given the intricate characteristics, advancement, inherent genetic composition, and diversity of cardiovascular diseases (CVDs). The appropriate utilization of artificial intelligence (AI) and machine learning (ML) methodologies can yield novel understandings of CVDs, enabling improved personalized treatments through predictive analysis and deep phenotyping. In this study, we proposed and employed a novel approach combining traditional statistics and a nexus of cutting-edge AI/ML techniques to identify significant biomarkers for our predictive engine by analyzing the complete transcriptome of CVD patients. After robust gene expression data pre-processing, we utilized three statistical tests (Pearson correlation, Chi-square test, and ANOVA) to assess the differences in transcriptomic expression and clinical characteristics between healthy individuals and CVD patients. Next, the Recursive Feature Elimination (RFE) classifier assigned rankings to transcriptomic features based on their relation to the case-control variable. The top ten percent of commonly observed significant biomarkers were evaluated using four unique ML classifiers (Random Forest, Support Vector Machine, Xtreme Gradient Boosting Decision Trees, and k-Nearest Neighbors). After optimizing hyperparameters, the ensembled models, which were implemented using a soft voting classifier, accurately differentiated between patients and healthy individuals. We have uncovered 18 transcriptomic biomarkers that are highly significant in the CVD population that were used to predict disease with up to 96% accuracy. Additionally, we cross-validated our results with clinical records collected from patients in our cohort. The identified biomarkers served as potential indicators for early detection of CVDs. With its successful implementation, our newly developed predictive engine provides a valuable framework for identifying patients with CVDs based on their biomarker profiles.

genomics↗

Integrated ACMG approved genes and ICD codes for the translational research and precision medicine

Timely understanding of biological secrets of complex diseases will ultimately benefit millions of individuals by reducing the high risks for mortality and improving the quality of life with personalized diagnoses and treatments. Due to the advancements in sequencing technologies and reduced cost, genomics data is developing at an unmatched pace and levels to foster translational research and precision medicine. Over ten million genomics datasets have been produced and publicly shared in the year 2022. Diverse and high-volume genomics and clinical data have the potential to broaden the scope of biological discoveries and insights by extracting, analyzing, and interpreting the hidden information. However, the current and still unresolved challenges include the integration of genomic profiles of the patients with their medical records. The disease definition in genomics medicine is simplified, when in the clinical world, diseases are classified, identified, and adopted with their International Classification of Diseases (ICD) codes, which are maintained by the World Health Organization (WHO). Several biological databases have been produced, which includes information about human genes and related diseases. However, still, there is no database exists, which can precisely link clinical codes with relevant genes and variants to support genomic and clinical data integration for clinical and translation medicine. In this project, we are focused on the development of an annotated gene-disease-code database, which is accessible through an online, cross-platform, and user-friendly application i.e., PAS-GDC. However, our scope is limited to the integration of ICD-9 and ICD-10 codes with the list of genes approved by the American College of Medical Genetics and Genomics (ACMG). Results include over seventeen thousand diseases and four thousand ICD codes, and over eleven thousand gene-disease-code combinations.

bioinformatics↗