Multiple instance learning on tile level-pathologyimages provides accurate and interpretableclassification for breast cancer molecular subtypes
Accurate breast cancer molecular subtyping is critical for treatment decisions, yet standard methods such as immunohistochemistry and gene expression profiling are costly and labor intensive. Deep learning classification approaches using Hematoxylin and Eosin-stained whole slide images are an active area of research. However, many existing methods rely on large, high-quality annotated datasets where tumor regions are manually outlined for segmentation. This process is costly, does not scale well, depends on expert pathologists, and may ignore relevant tumor microenvironment features or reflect subjective labelling decisions. Here, we present an annotation-free, weakly supervised pipeline and web-based tool for breast cancer molecular subtyping using a computational pathology foundation model. A total of 1433 WSIs from three public cohorts (TCGA-BRCA, CPTAC-BRCA, and the Warwick HER2 cohort) were tiled into 224x224 patches without overlap at 20x magnification. Tile-level embeddings were extracted with a foundation model, and slide-level representations were obtained by mean pooling. We evaluated one-vs-rest classifiers including cosine similarity, logistic regression, and attention-based multiple instance learning. On a held-out test set of 287 WSIs, calibrated logistic regression achieved a macro F1 score of 0.75 using slide embeddings, while attention-based MIL reached 0.83 using tile embeddings. Luminal A and Basal subtypes were predicted reliably, whereas Luminal B remained challenging. Novel attention and probability heatmaps highlights spatial regions most informative for predictions, supporting qualitative interpretability. These results demonstrate accurate and interpretable breast cancer subtyping without tumor annotations, and we provide a web server to support pathology diagnostics.