Search bioRxiv⌕ Search

Biology subjects

Jamal, T. B.

Publications and source records attributed to Jamal, T. B..

4 recordsLinked to original sources

A Non-invasive Detection of Parkinson's Disease using PitArray: An Integrative Meta-Analysis and Machine Learning Approach

Parkinsons disease (PD) is a progressive neurodegenerative disorder affecting the central nervous system, often diagnosed in its advanced stages due to the absence of sensitive biomarkers. With this objective in mind, our study conducted a comprehensive analysis of differentially expressed genes (DEGs) sourced from blood-based microarray datasets to uncover potential biomarkers and developed a machine learning based classifier to conduct two step validations. By analyzing gene expression of three projects, we identified 678 DEGs, consisting of 337 genes showing upregulation and 341 genes presenting downregulation. Additionally, insights from functional enrichment and the protein-protein network analysis indicate that HLA-F, IRF-1, and RPS28 have the potential to serve as biomarkers for diagnosing PD. Simultaneously, we employed feature selection techniques such as Least Absolute Shrinkage and Selection Operator with Cross Validation (LassoCV) followed by Recursive Feature Elimination with Cross Validation (REFCV) to filter our initial dataset of 13,249 genes down to 43 genes, which were subsequently used to train the machine learning-based classifier models. These 43 genes formed the basis for training and testing various machine learning models, including logistic regression, random forest, naive Bayes, k-nearest neighbors, support vector machine, and deep learning based artificial neural networks. Our models demonstrated robust performance, with Support Vector Machine outperforming others by 0.65 accuracy (95%CI: 0.58-0.66), 0.70 AUC-ROC (95%CI: 0.70-0.71) and 0.35 MCC (95%CI: 0.34-0.39). The model was implemented to develop the PitArray tool for non-invasive detection of PD from blood. PitArray is available at: https://github.com/Arittra95/PitArray. Key PointsO_LIHLA-F, IRF-1, and RPS28 were identified as potential biomarkers for Parkinsons disease diagnosis. C_LIO_LISeveral sophisticated feature selection methods recognized 43 genes which were then used to build a machine learning model. C_LIO_LIA Support Vector Machine based tool named PitArray was developed which could distinguish Parkinsons disease patients from healthy people based on blood transcriptome data. C_LI

bioinformatics↗

An Integrated Comparative Genomics, Subtractive Proteomics and Immunoinformatics Framework for the Rational Design of a Pan-Salmonella Multi-Epitope Vaccine

Salmonella infections are a global public health issue due to the high cost of illness surveillance, prevention, and treatment. In this study, we explored the core proteome in Salmonella to design a multi-epitope vaccine through Subtractive Proteomics and immunoinformatics approaches. A total of 2395 core proteins presents in 30 different strains of Salmonella (reference strain-NZ CP014051) were curated. Utilizing the subtractive proteomics approach on the Salmonella core proteome, Curlin major subunit A (CsgA) was selected as the vaccine candidate. csgA is a conserved gene that is related with biofilm formation. Immunodominant B and T cell epitopes from CsgA were predicted using numerous immunoinformatics tools. T lymphocyte epitopes had adequate population coverage and their corresponding MHC alleles showed significant binding scores after peptide-protein based molecular docking. Afterward, a multiepitope vaccine was constructed with peptide linkers and Human Beta Defensin-2 (as an adjuvant). The vaccine was found to be highly antigenic, non-toxic, non-allergic, and had physicochemical properties. Additionally, Molecular Dynamics Simulation and Immune Simulation demonstrated that the vaccine can bind with Toll Like Receptor 4 and elicit robust immune response. Using in vitro, in vivo, and clinical trials, our results would yield a Pan-Salmonella vaccine that will provide protection against various Salmonella species.

bioinformatics↗

AITeQ: A machine learning framework for Alzheimer's prediction using a distinctive 5-gene signature

Neurodegenerative diseases, such as Alzheimers disease, pose a significant global health challenge with their complex etiology and elusive biomarkers. In this study, we developed the Alzheimers Identification Tool using RNA-Seq (AITeQ), a machine learning model based on an optimized random forest algorithm for identification of Alzheimers from RNA-Seq data. Analysis of RNA-Seq data from 433 individuals, including 293 Alzheimers patients and 140 controls led to the discovery of 47,929 differentially expressed genes. This was followed by a machine learning protocol involving feature selection, model training, performance evaluation, and hyperparameter tuning. The feature selection process undertaken in this study, employing a combination of 4 different methodologies, culminated in the identification of a compact yet impactful set of 5 genes. Ten diverse machine learning models were trained and tested using these 5 genes (ITGA10, CXCR4, ADCYAP1, SLC6A12, VGF). Performance metrics, including precision, recall, F1-score, accuracy, receiver operating characteristic area under the curve, and confusion matrices, were assessed before and after hyperparameter tuning. Overall, the random forest model with optimized hyperparameters was identified as the best and was used to develop AITeQ. AITeQ is available at: https://github.com/ishtiaque-ahammad/AITeQ Key PointsO_LIA set of 5 genes (ITGA10, CXCR4, ADCYAP1, SLC6A12, VGF) were identified following differential gene expression and feature importance analysis. C_LIO_LITen diverse machine learning algorithms were trained and tested using the gene expression patterns of the identified 5 genes. The random forest algorithm with customized hyperparameters was found to be the best-performing model for differentiating Alzheimers disease samples from control. C_LIO_LIAITeQ, a user-friendly, reliable, and accurate machine learning framework for Alzheimers disease prediction was developed based on the 5 gene signature. C_LI

bioinformatics↗

Shotgun metagenomics unravels higher antibiotic resistome profile in Bangladeshi gut microbiome

Antibiotic resistance management is a challenging task in Low and Middle-Income Countries (LMICs) such as Bangladesh. Improper regulation and uncontrolled spreading of Antibiotic Resistant Genes (ARGs) from LIMCs pose a great threat to global public health. The human gut microbiome is a massive reservoir of Antibiotic Resistant Genes (ARGs). In this study, we unraveled the ARGs in the gut microbiome of the Bangladeshi population and compared them with several other countries around the world. Here, 31 fecal samples from different ethnic groups living in Bangladesh namely Bengali (n=9), Chakma (n=6), Khyang (n=5), Marma (n=6), and Tripura (n=5) were collected. Shotgun metagenomic sequencing method was implemented for revealing the ARGs. The resistome profiling was executed on three levels-the total microbiome, the plasmidome, and the virome. In all three levels, samples from Bangladeshi cohorts showed higher ARG profiles compared to foreign samples. On average, the number of ARGs in the Bangladeshi samples ranged between 75.11 and 88. Among them, class C beta-lactamases, quinolone resistance genes, and tetracycline efflux pumps were relatively more abundant. Additionally, the MexPQ-OpmE drug resistance pathway was found to be more prevalent. Findings from our study suggest that the spread of antibiotic resistance within the Bangladeshi population is being facilitated by the gut microbiome especially via the mobilome. Therefore, strict regulation on antibiotic usage is necessary to halt the spread of ARGs.

genomics↗