Search bioRxiv⌕ Search

Biology subjects

Shekhawat, P.

Publications and source records attributed to Shekhawat, P..

1 recordsLinked to original sources

AllerStack: Predicting Allergenic Proteins with a Stacked Ensemble Approach

Accurate prediction of protein allergenicity is essential for ensuring food and drug safety. While machine learning and deep learning models have been explored for this task, limitations remain in dataset scale, feature representation, and model architecture. Here, we introduce AllerStack, a two-stage stacked ensemble model integrating handcrafted and ESM2-based learned features for allergenicity classification. The model was developed using a balanced dataset comprising 11,930 allergenic and 11,930 non-allergenic proteins. We extracted amino acid composition, dipeptide composition, and physicochemical features using the Biopython library, alongside contextual embeddings from the pre-trained ESM2 protein language model. Diverse classifiers (QDA, SVM, KNN, and ANN) were trained separately on these features in the base layer. Their predictions were used as input to a meta-classifier based on XGBoost. AllerStack achieved high predictive performance with 96.87% accuracy, 96.86% F1-score, 93.75% Matthews correlation coefficient (MCC), and an AUC of 0.99. A publicly accessible web server (https://cosylab.iiitd.edu.in/allerstack/) enables real-time allergenicity prediction from protein sequences. AllerStack provides a robust, interpretable, and user-friendly platform for allergen detection in computational biology. HighlightsO_LIAn extensive dataset of 23,860 proteins. C_LIO_LICombines handcrafted and ESM2-derived features. C_LIO_LIStacked ensemble model with an accuracy of 96.87%. C_LIO_LISHAP-based model interpretability at the model and feature level. C_LIO_LIPublic web server, AllerStack (https://cosylab.iiitd.edu.in/allerstack/). C_LI

bioinformatics↗