Search bioRxiv⌕ Search

Biology subjects

Fisher, N. C.

Publications and source records attributed to Fisher, N. C..

5 recordsLinked to original sources

Library size can undermine accurate molecular and phenotypic subtyping in spatial transcriptomics data.

In an era where transcriptomics-based subtyping, phenotyping and mechanistic understanding is increasingly being driven by state-of-the-art spatially resolved transcriptomic (ST) technologies, it is imperative that researchers, journals, and funders do all they can to ensure that as a community we are interpreting these exciting data as accurately as possible with awareness of their limitations. In this short report, we highlight one potential bias in ST data that could undermine accurate interpretation of transcriptional signatures, providing the field with an opportunity to identify and avoid this issue prior to release of new mechanistic findings. This issue is particularly relevant for platforms that produce some of the most granular and high-resolution spatial information at single cell (and sub-cellular) resolution, with the compromise of a reduced transcriptome panel of genes (Figure 1A). O_FIG O_LINKSMALLFIG WIDTH=141 HEIGHT=200 SRC="FIGDIR/small/602370v1_fig1.gif" ALT="Figure 1"> View larger version (42K): org.highwire.dtl.DTLVardef@1219fb3org.highwire.dtl.DTLVardef@7bd4a4org.highwire.dtl.DTLVardef@1c55c42org.highwire.dtl.DTLVardef@2bf924_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 1.C_FLOATNO Visualisation of the iCMS3 up signature A: Schematic overview of enrichment using a full gene panel versus a reduced panel. B: Overlap between the genes represented on the CosMx and Xenium gene panels. C: Correlation between the single sample scores of the full iCMS3 up signature (74 genes) and the 15 genes from the signature represented on the CosMx platform. Two samples are highlighted which have a similar iCMS3 up enrichment for the full signature (CRC-JSC-S06: 4535.746; SMC16: 4265.263) but extreme enrichments for the signature composed only of the genes present on the CosMx array (CRC-JSC-S06: 301.0396; SMC16: 11430.593). Median (4047.994) shown by red line. D: Visualisation of the rank of each sample across for the full (CRC-JSC-S06: position 25872/44458; SMC16: position 23907/44458) and CosMx (CRC-JSC-S06: position 513/44458; SMC16: position 44241/44458) signature enrichment. E: Heatmap showing the relative enrichment of each gene with the iCMS3 up signature for each sample, with the genes present on the CosMx array in red. F: Subset of samples (n=628) +/-1% of the median (4007.514 - 4088.474), with the top 100 and bottom 100 samples for CosMx enrichment in red. G: Heatmap of the top 100 and bottom 100 samples (shown in red in F) for CosMx enrichment with a full iCMS3 up enrichment around the median. The samples are arranged by the rank of each sample for CosMx signature enrichment, with the sum of the genes within the iCMS3 up signature present on the array (CosMx [n=16]) and those not present on the array (Non CosMx [n=58]) overlaid as barplots. C_FIG

cancer biology↗

Evaluation of Gene Set Enrichment Analysis (GSEA) tools highlights the value of single sample approaches over pairwise for robust biological discovery.

BackgroundGene set enrichment analysis (GSEA) tools can be used to identify biological insights from transcriptional datasets and have become an integral analysis within gene expression-based cancer studies. Over the years, additional methods of GSEA-based tools have been developed, providing the field with an ever-expanding range of options to choose from. Although several studies have compared the statistical performance of these tools, the downstream biological implications that arise when choosing between the range of pairwise or single sample forms of GSEA methods remain understudied. MethodsIn this study, we compare the statistical and biological interpretation of results obtained when using a variety of pre-ranking methods and options for pairwise GSEA and fast GSEA (fGSEA), alongside single sample GSEA (ssGSEA) and gene set variation analysis (GSVA). These analyses are applied to a well-established cohort of n=215 colon tumour samples, using the clinical feature of cancer recurrence status, non-relapse (NR) and relapse (R), as an initial exemplar, in conjunction with the Molecular Signatures Database "Hallmark" gene sets. ResultsDespite minor fluctuations in statistical performance, pairwise analysis revealed remarkably similar results when deployed using a range of gene pre-ranking methods or across a range of choices of GSEA versus fGSEA, with the same well-established prognostic signatures being consistently returned as significantly associated with relapse status. In contrast, when the same statistically significant signatures, such as Interferon Gamma Response, were assessed using ssGSEA and GSVA approaches, there was a complete absence of biological distinction between these groups (NR and R). ConclusionsData presented here highlights how pairwise methods can overgeneralise biological enrichment within a group, assigning strong statistical significance to gene sets that may be inadvertently interpreted as equating to distinct biology. Importantly, single sample approaches allow users to clearly visualise and interpret statistical significance alongside biological distinction between samples within groups-of-interest; thus, providing a more robust and reliable basis for discovery research.

cancer biology↗

MmCMS: Mouse models' Consensus Molecular Subtypes of colorectal cancer

BACKGROUNDColorectal cancer (CRC) primary tumours are molecularly classified into four consensus molecular subtypes (CMS1-4). Genetically engineered mouse models aim to faithfully mimic the complexity of human cancers and, when appropriately aligned, represent ideal pre-clinical systems to test new drug treatments. Despite its importance, dual-species classification has been limited by the lack of a reliable approach. Here we utilise, develop and test a set of options for human-to-mouse CMS classifications of CRC tissue. METHODSUsing transcriptional data from established collections of CRC tumours, including human (TCGA cohort; n=577) and mouse (n=57 across n=8 genotypes) tumours with combinations of random forest and nearest template prediction algorithms, alongside gene ontology collections, we comprehensively assess the performance of a suite of new dual-species classifiers. RESULTSWe developed three approaches: MmCMS-A; a gene-level classifier, MmCMS-B; an ontology-level approach and MmCMS-C; a combined pathway system encompassing multiple biological and histological signalling cascades. Although all options could identify tumours associated with stromal-rich CMS4-like biology, MmCMS-A was unable to accurately classify the biology underpinning epithelial-like subtypes (CMS2/3) in mouse tumours. CONCLUSIONSWhen applying human-based transcriptional classifiers to mouse tumour data, a pathway-level classifier, rather than an individual gene-level system, is optimal. Our R package with three options helps researchers select suitable mouse models of human CRC subtype for their experimental testing.

cancer biology↗

Biological misinterpretation of transcriptional signatures in tumour samples can unknowingly undermine mechanistic understanding and faithful alignment with preclinical data

Precise mechanism-based gene expression signatures (GESs) have been developed in appropriate in vitro and in vivo model systems, to identify important cancer-related signalling processes. However, some GESs originally developed to represent specific disease processes, primarily with an epithelial cell focus, are being applied to heterogeneous tumour samples where the expression of the genes in the signature may no longer be epithelial-specific. Therefore, unknowingly, even small changes in tumour stroma percentage can directly influence GESs, undermining the intended mechanistic signalling. Using colorectal cancer as an exemplar, we deployed numerous orthogonal profiling methodologies, including laser capture microdissection, flow cytometry, bulk and multiregional biopsy clinical samples, single cell RNAseq and finally spatial transcriptomics, to perform a comprehensive assessment of the potential for the most widely-used GESs to be influenced, or confounded, by stromal content in tumour tissue. To complement this work, we generated a freely-available resource, ConfoundR; https://confoundr.qub.ac.uk/, that enables users to test the extent of stromal influence on an unlimited number of the genes/signatures simultaneously across colorectal, breast, pancreatic, ovarian and prostate cancer datasets. Findings presented here demonstrate the clear potential for misinterpretation of the meaning of GESs, due to widespread stromal influences, which in-turn can undermine faithful alignment between clinical samples and preclinical data/models, particularly cell lines and organoids, or tumour models not fully recapitulating the stromal and immune microenvironment. As such, efforts to faithfully align preclinical models of disease using phenotypically-designed GESs must ensure that the signatures themselves remain representative of the same biology when applied to clinical samples.

cancer biology↗

Development of a semi-automated method for tumor budding assessment in colorectal cancer and comparison with manual methods

Tumor budding is an established prognostic feature in multiple cancers but routine assessment has not yet been incorporated into clinical pathology practice. Recent efforts to standardize and automate assessment have shifted away from haematoxylin and eosin (H&E)-stained images towards cytokeratin (CK) immunohistochemistry. In this study, we compare established manual H&E and cytokeratin budding assessment methods with a new, semi-automated approach built within the QuPath open-source software. We applied our method to tissue cores from the advancing tumor edge in a cohort of stage II/III colon cancers (n=186). The total number of buds detected by each method, over the 186 TMA cores, were as follows; manual H&E (n=503), manual CK (n=2290) and semi-automated (n=5138). More than four times the number of buds were detected using CK compared to H&E. A total of 1734 individual buds were identified both using manual assessment and semi-automated detection on CK images, representing 75.7% of the total buds identified manually (n=2290) and 33.7% of the total buds detected using our proposed semi-automated method (n=5138). Higher bud scores by the semi-automated method were due to any discrete area of CK immunopositivity within an accepted area range being identified as a bud, regardless of shape or crispness of definition, and to inclusion of tumor cell clusters within glandular lumina ("luminal pseudobuds"). Although absolute numbers differed, semi-automated and manual bud counts were strongly correlated across cores ({rho}=0.81, p<0.0001). Despite the random, rather than "hotspot", nature of tumor core sampling, all methods of budding assessment demonstrated poorer survival associated with higher budding scores. In conclusion, we present a new QuPath-based approach to tumor budding assessment, which compares favorably to current established methods and offers a freely-available, rapid and transparent tool that is also applicable to whole slide images.

pathology↗