A robust unsupervised clustering approach for high-dimensional biological imaging data reveals shared drug-induced morphological signatures
Modern biology increasingly relies on large-scale screening to generate high dimensional datasets with potential to accelerate discovery. However, analysing these complex datasets remains challenging, particularly in applications where the underlying structure and groupings are unknown, and high dimensionality introduces noise and artifacts that make follow up studies difficult to prioritise. Here, we present an unsupervised consensus clustering tool that quantifies biologically meaningful patterns based on multi-scale data organisation to guide decision-making in high-throughput screening. Using large-scale drug screening data in cancer cell lines and bacterium model, we demonstrate its ability to use diverse data inputs to prioritize robust drug clusters with shared biological mechanisms and conserved drug responses. This method addresses key limitations associated with prioritising robust, actionable hits from scalable screening data.