Search bioRxiv⌕ Search

Biology subjects

Leh, S.

Publications and source records attributed to Leh, S..

2 recordsLinked to original sources

StainStyleSampler: Clustering-based sampling of whole slide image appearances

The appearance of whole slide biopsy images is greatly affected by various factors such as laboratory procedures or the choice of digital slide scanners. The resulting variations in image styles within and across batches of histological images represent one of the major obstacles to the development of generalizable machine learning algorithms. To overcome this challenge, a lot of research has focused on stain normalization and stain augmentation techniques. While such approaches provide effective strategies to reduce stain variation or increase stain invariance, respectively, they typically involve only limited modeling or sampling of the underlying stain style distribution. Tools for a streamlined sampling of different aspects of such a distribution, which would be crucial e.g. for explicitly evaluating machine learning robustness across or with respect to major stain styles, remain largely missing. Here, we present the StainStyleSampler, a toolkit for (i) the exploration and modeling of stain style variations, and (ii) the automated sampling of images or styles capturing the core components of this variation. The tool enables the extraction of various color features and deconvolved stain components, visualization of such features directly or after dimensionality reduction, modeling of style distributions using binning, clustering, and density mapping, and automated sampling of the most representative reference images. We believe that this software will equip pathologists and computer-scientists with a more versatile set of tools that can aid substantially in both the exploration and sampling of stain variation across whole slide images.

pathology↗

Unsupervised learning for labeling global glomerulosclerosis

Current deep learning models for classifying glomeruli in nephropathology are trained almost exclusively in a supervised manner, requiring expert-labeled images. Very little is known about the potential for unsupervised learning to overcome this bottleneck. To address this open question in a proof-of-concept, the project focused on the most fundamental classification task: globally sclerosed versus non-globally sclerosed glomeruli. The performance of clustering between the two classes was extensively studied across a variety of labeled datasets with diverse compositions and histological stains, and across the feature embeddings produced by 34 different pre-trained CNN models. As demonstrated by the study, clustering of globally and non-globally sclerosed glomeruli is generally highly feasible, yielding accuracies of over 95% in most datasets. Further work will be required to expand these experiments towards the clustering of additional glomerular lesion categories. We are convinced that these efforts (i) will open up opportunities for semi-automatic labeling approaches, thus alleviating the need for labor-intensive manual labeling, and (ii) illustrate that glomerular classification models can potentially be trained even in the absence of expert-derived class labels.

pathology↗