Search bioRxiv⌕ Search

Biology subjects

Taherian Fard, A.

Publications and source records attributed to Taherian Fard, A..

3 recordsLinked to original sources

How does data structure impact cell-cell similarity? Evaluating the influence of structural properties on proximity metric performance in single cell RNA-seq data

Accurately identifying cell populations is paramount to the quality of downstream analyses and overall interpretations of single-cell RNA-seq (scRNA-seq) datasets but remains a challenge. The quality of single-cell clustering depends on the proximity metric used to generate cell-to-cell distances. Accordingly, proximity metrics have been benchmarked for scRNA-seq clustering, typically with results averaged across datasets to identify a highest performing metric. However, the best-performing metric varies between studies, with the performance differing significantly between datasets. This suggests that the unique structural properties of a scRNA-seq dataset, specific to the biological system under study, has a substantial impact on proximity metric performance. Previous benchmarking studies have omitted to factor the structural properties into their evaluations. To address this gap, we developed a framework for the in-depth evaluation of the performance of 17 proximity metrics with respect to core structural properties of scRNA-seq data, including sparsity, dimensionality, cell population distribution and rarity. We find that clustering performance can be improved substantially by the selection of an appropriate proximity metric and neighbourhood size for the structural properties of a dataset, in addition to performing suitable pre-processing and dimensionality reduction. Furthermore, popular metrics such as Euclidean and Manhattan distance performed poorly in comparison to several lessor applied metrics, suggesting the default metric for many scRNA-seq methods should be re-evaluated. Our findings highlight the critical nature of tailoring scRNA-seq analyses pipelines to the system under study and provide practical guidance for researchers looking to optimise cell similarity search for the structural properties of their own data.

bioinformatics↗

scShapes: A statistical framework for identifying distribution shapes in single-cell RNA-sequencing data

BackgroundSingle cell RNA sequencing (scRNA-seq) methods have been advantageous for quantifying cell-to-cell variation by profiling the transcriptomes of individual cells. For scRNA-seq data, variability in gene expression reflects the degree of variation in gene expression from one cell to another. Analyses that focus on cell-cell variability therefore are useful for going beyond changes based on average expression and instead, identifying genes with homogenous expression versus those that vary widely from cell to cell. ResultsWe present a novel statistical framework scShapes for identifying differential distributions in single-cell RNA-sequencing data using generalized linear models. Most approaches for differential gene expression detect shifts in the mean value. However, as single cell data are driven by over-dispersion and dropouts, moving beyond means and using distributions that can handle excess zeros is critical. scShapes quantifies gene-specific cell-to-cell variability by testing for differences in the expression distribution while flexibly adjusting for covariates if required. We demonstrate that scShapes identifies subtle variations that are independent of altered mean expression and detects biologically-relevant genes that were not discovered through standard approaches. ConclusionsThis analysis also draws attention to genes that switch distribution shapes from a unimodal distribution to a zero-inflated distribution and raises open questions about the plausible biological mechanisms that may give rise to this, such as transcriptional bursting. Overall, the results from scShapes helps to expand our understanding of the role that gene expression plays in the transcriptional regulation of a specific perturbation or cellular phenotype. Our framework scShapes is incorporated into Bioconductor R package (https://github.com/Malindrie/scShapes).

bioinformatics↗

Deconstructing replicative senescence heterogeneity of human mesenchymal stem cells at single cell resolution reveals therapeutically targetable senescent cell sub-populations

Cellular senescence is characterised by a state of permanent cell cycle arrest. It is accompanied by often variable release of the so-called senescence-associated secretory phenotype (SASP) factors, and occurs in response to a variety of triggers such as persistent DNA damage, telomere dysfunction, or oncogene activation. While cellular senescence is a recognised driver of organismal ageing, the extent of heterogeneity within and between different senescent cell populations remains largely unclear. Elucidating the drivers and extent of variability in cellular senescence states is important for discovering novel targeted seno-therapeutics and for overcoming cell expansion constraints in the cell therapy industry. Here we combine cell biological and single cell RNA-sequencing approaches to investigate heterogeneity of replicative senescence in human ESC-derived mesenchymal stem cells (esMSCs) as MSCs are the cell type of choice for the majority of current stem cell therapies and senescence of MSC is a recognized driver of organismal ageing. Our data identify three senescent subpopulations in the senescing esMSC population that differ in SASP, oncogene expression, and escape from senescence. Uncovering and defining this heterogeneity of senescence states in cultured human esMSCs allowed us to identify potential drug targets that may delay the emergence of senescent MSCs in vitro and perhaps in vivo in the future.

bioinformatics↗