bioRxiv · 10.1101/2024.02.07.579420
Discovering and overcoming the bias in neoantigen identification by unified machine learning models
Abstract
Neoantigens play a crucial role in tumor immune process and precisely identifying them can greatly contribute to tumor immunotherapy design. There are three main steps in the neoantigen immune process, i.e., binding with MHCs, extracellular presentation, and immunogenicity induction. Various computational methods have been developed, but the overall accuracy of neoantigen identification remains relatively low. Here, we established a unified transformer-based framework ImmuBPI that comprised three tasks. Cross-task model interpretation discovered a counterfactual pattern learned by the immunogenicity prediction model overlooked in previous studies. We demonstrated that this model bias arose from training data imbalance and the neglect of bias control would not only mislead the models but also hinder existing benchmarks from fair evaluation. We further designed a mutual information-based debiasing strategy that showed preliminary potential in alleviating the bias. We believe these observations will provide insightful perspectives for future neoantigen prediction by high-lighting the necessity of bias control.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zhang, Z., Wu, W., Wei, L., Wang, X.. 2024-02-08. Discovering and overcoming the bias in neoantigen identification by unified machine learning models. https://doi.org/10.1101/2024.02.07.579420
Cite the original work for its findings. Save a collection to share your selection of sources.