Search bioRxiv⌕ Search

Biology subjects

Selega, A.

Publications and source records attributed to Selega, A..

3 recordsLinked to original sources

Segmentation error aware clustering for highly multiplexed imaging

Spatial protein expression technologies can map cellular content and organization by simultaneously quantifying the expression of >40 proteins at subcellular resolution within intact tissue sections and cell lines. However, necessary image segmentation to single cells is challenging and error prone, easily confounding the interpretation of cellular phenotypes and cell clusters. To address these limitations, we present STARLING, a novel probabilistic machine learning model designed to quantify cell populations from spatial protein expression data while accounting for segmentation errors. To evaluate performance we developed a comprehensive benchmarking workflow by generating highly multiplexed imaging data of cell line pellet standards with controlled cell content and marker expression and additionally established a novel score to quantify the biological plausibility of discovered cellular phenotypes on patient derived tissue sections. Moreover, we generate spatial expression data of the human tonsil - a densely packed tissue prone to segmentation errors - and demonstrate cellular states captured by STARLING identify known cell types not visible with other methods and enable quantification of intra- and inter- individual heterogeneity. STARLING is available at https://github.com/camlab-bioml/starling.

bioinformatics↗

Beyond benchmarking: towards predictive models of dataset-specific single-cell RNA-seq pipeline performance

The advent of single-cell RNA-sequencing (scRNA-seq) has driven significant computational methods development for all steps in the scRNA-seq data analysis pipeline, including filtering, normalization, and clustering. The large number of methods and their resulting parameter combinations has created a combinatorial set of possible pipelines to analyze scRNA-seq data, which leads to the obvious question: which is best? Several benchmarking studies have sought to compare methods to answer this, but frequently find variable performance depending on dataset and pipeline characteristics. Alternatively, the large number of publicly available scRNA-seq datasets along with advances in supervised machine learning raise a tantalizing possibility: could the optimal pipeline be predicted for a given dataset? Here we begin to answer this question by applying 288 scRNA-seq analysis pipelines to 86 datasets and quantifying pipeline success via a range of measures evaluating cluster purity and biological plausibility. We build supervised machine learning models to predict pipeline success given a range of dataset and pipeline characteristics. We find both that prediction performance is significantly better than random and that in many cases pipelines predicted to perform well provide clustering outputs similar to expert-annotated cell type labels. Finally, we identify characteristics of scRNA-seq datasets that correlate with strong prediction performance that could guide when such prediction models may be useful.

bioinformatics↗

Multi-objective Bayesian Optimization with Heuristic Objectives for Biomedical and Molecular Data Analysis Workflows

Many practical applications require optimization of multiple, computationally expensive, and possibly competing objectives that are well-suited for multi-objective Bayesian optimization (MOBO) procedures. However, for many types of biomedical data, measures of data analysis workflow success are often heuristic and therefore it is not known a priori which objectives are useful. Thus, MOBO methods that return the full Pareto front may be suboptimal in these cases. Here we propose a novel MOBO method that adaptively updates the scalarization function using properties of the posterior of a multi-output Gaussian process surrogate function. This approach selects useful objectives based on a flexible set of desirable criteria, allowing the functional form of each objective to guide optimization. We demonstrate the qualitative behaviour of our method on toy data and perform proof-of-concept analyses of single-cell RNA sequencing and highly multiplexed imaging datasets.

bioinformatics↗