Search bioRxiv⌕ Search

Biology subjects

Peres, L. C.

Publications and source records attributed to Peres, L. C..

6 recordsLinked to original sources

Functional Data Analysis of Spatial Clustering Identifies Prognostic T Cell Patterns in Ovarian Cancer

Spatial proteomic imaging technologies enable the simultaneous assessment of immune cell abundance and spatial organization within the tumor microenvironment. Spatial clustering is commonly summarized using measures such as Ripleys K or nearestOneighbor G-functions at a fixed radius. However, these approaches depend on scale selection and may obscure biologically relevant patterns occurring across spatial ranges. We propose a functional data analysis (FDA) framework to model spatial clustering trajectories derived across a continuum of radii. Functional principal component analysis (FPCA) was used to summarize dominant modes of spatial variation, and resulting scores were incorporated into Cox proportional hazards models as both main effects and interaction with immune cell abundance. The approach was applied to multiplex immunofluorescence data from five ovarian cancer studies, comprising 773 highOgrade ovarian serous tumors. Analyses focused on CD3+ and CD8+ T cell populations within the tumor compartment of the tissue, adjusting for age at diagnosis and cancer stage, with study-specific estimates combined using random-effects meta-analysis. Higher abundance of both T cells and CD8+ T cells was consistently associated with improved overall survival. Beyond abundance, spatial features captured by the leading functional principal component were independently associated with survival, particularly for CD8+ T cells. Interaction models further showed that the prognostic effect of immune infiltration depended on spatial clustering, with tumors characterized by high abundance and low spatial clustering exhibiting the most favorable outcomes. These findings indicate that spatial organization provides complementary prognostic information beyond abundance alone and suggests that more diffuse immune infiltration may reflect more effective anti-tumor activity in ovarian cancer. Overall, FDA offers a flexible and interpretable framework for modeling spatial clustering across scales and identifying prognostic spatial features not captured by fixed-radius or distance analyses.

cancer biology↗

Transcriptomic subtypes in high-grade serous ovarian cancer are driven by tumor cellular composition

High-grade serous ovarian carcinoma (HGSC) is an aggressive malignancy for which bulk transcriptomic subtypes are used to stratify tumors, interpret biology, and guide biomarker development. The four TCGA-derived subtypes, mesenchymal (C1.MES), immunoreactive (C2.IMM), proliferative (C5.PRO), and differentiated (C4.DIF), are consistently observed across cohorts. However, despite their prominence, these subtypes have not translated into therapeutic utility, and their biological basis remains unresolved. Here, we show that HGSC transcriptomic subtypes are largely determined by tumor cellular composition rather than intrinsic malignant transcriptional programs. By integrating controlled single-cell-derived pseudobulk simulations with deconvolution-based analysis of 1,834 primary HGSC tumors across RNA-seq and microarray cohorts, we demonstrate that subtype probabilities align along a composition-driven axis of stromal and immune variation. Cellular composition alone predicted subtype labels with high accuracy (ROC-AUC = 0.81-0.95) and explained a substantial fraction of subtype-associated transcriptomic variation, with the mesenchymal (C1.MES) subtype representing the most robust and reproducible example of composition-driven signal. Although a secondary, composition-independent expression signal is detectable, it does not define the dominant structure of subtype classification. These findings redefine HGSC transcriptomic subtypes as features of the tumor ecosystem rather than discrete malignant states. This reinterpretation has immediate implications for studies that use subtype labels to infer tumor-intrinsic biology and provides a generalizable framework for separating composition-driven and intrinsic signals in bulk tumor data. Significance StatementHGSC transcriptomic subtypes lack consistent clinical utility and remain biologically ambiguous. We show subtype assignments are largely driven by tumor cellular composition, and less so by distinct intrinsic tumor states.

cancer biology↗

Deconvolved tumor adipocyte proportions and high grade serous ovarian carcinoma survival

BackgroundSingle-cell-based analyses of high-grade serous ovarian carcinoma (HGSOC) survival have largely ignored adipocytes, which are fragile and under-represented in single-cell references. Adipocytes are known active components of the tumor microenvironment in many cancers, and HGSOC tumors frequently metastasize to the omentum, a lining of adipose tissue. MethodsWe created a composite reference that combines single-nucleus adipose profiles with published HGSOC single-cell data to deconvolve 588 bulk RNA-seq tumours from the Schildkraut cohorts. We used stage-stratified Cox models to quantify the association between intratumoural adipocyte fractions and overall survival while adjusting for age, body mass index (BMI), race, and residual disease. We also evaluated associations with deconvolved immune, stromal, and epithelial cell groups. ResultsA 10% increase in estimated tumor adipocyte content was associated with a 41% increase in the hazard of death (HR = 1.41, 95% CI 1.18-1.70, p = 0.0002) after adjusting for age, BMI and race (n=566). A 10% increase in immune cell proportion was associated with favorable survival (HR = 0.82, 95% CI 0.69-0.97, p = 0.024). Stromal and epithelial macro-fractions were not associated with survival. Associations with adipocyte and immune cell type proportions were unchanged in models additionally controlling the other cell type proportions. Results were similar after additionally adjusting for residual disease after debulking surgery. ConclusionsAdipocytes may be a tumor-intrinsic factor associated with adverse outcomes in HGSOC. Quantifying adipocyte burden using bulk RNA-seq could enhance risk stratification and guide the development of adipocyte-targeted therapies.

genomics↗

Exact Expectation of Complete Spatial Randomness for Nearest Neighbor G(r): A Scalable Alternative to Permutations

Spatial analysis is becoming increasingly important for studies, from epidemiology to tissue biology, as technologies advance and experimental costs decrease. However, the widespread use of spatial metrics such as Nearest Neighbor G(r) is affected by the fact that biological systems rarely satisfy the assumption of stationarity, which is required to appropriately use theoretical complete spatial randomness (CSR) measures. As a result researchers often use computationally expensive permutations to empirically estimate CSR for subsets of points or cells. Here, we present closed form analytical solutions for both the mean and variance of the sample-specific CSR for Nearest Neighbor G(r) to allow for fast and reproducible calculation without permutations. Using a multiplex immunofluorescence sample of clear cell renal cell carcinoma, we show that the theoretical G(r) for cytotoxic T cells overestimates CSR at low radii (due to spatial constraints between cells) while drastically underestimating CSR at radii between 20 and 90 pixels. In a simulated sample of 30 points, our analytical solution for the mean is identical to the average of G(r) measured on all 142,506 unique combinations of 5 marked (or positive) points. On the real clear cell renal cell carcinoma sample, our exact CSR is similar in speed to estimating CSR with 1000 permutations while our optimized Rcpp implementation is [~]30x faster and consuming [~]20x less memory than 1000 permutations. This permutation-free approach dramatically enhances computational efficiency and reproducibility, enabling scalable and reproducible analysis for studies in epidemiology, multiplex immunofluorescence, spatial transcriptomics, and related fields where accurate, sample-specific null expectations are important for comparisons.

bioinformatics↗

scSpatialSIM: a simulator of spatial single-cell molecular data

BackgroundSpatial molecular data is increasingly being generated in biological tissue studies to increase our understanding of cell infiltration and spatial architecture of tissues. Examples of technologies used to study the spatial contexture of tissues are single-cell protein expression assays and spatial transcriptomics. The increased use of spatial biology technologies has also resulted in an increase in the development of statistical methods to describe the spatial landscape in tissues. Due to the lack of consensus on "gold standard" statistical approaches for assessing the spatial contexture of tissues, we created an R package, scSpatialSIM, to assess different statistical and bioinformatic methods. scSpatialSIM allows users to simulate single-cell molecular data to mimic real tissues at scale, clustering of cell types, and co-clustering / co-localization of two or more cell types. scSpatialSIM also contains functions that give users the ability to simulate quantitative distributions for positive and negative cells (e.g., gene expression, fluorescence intensity). ResultsWe demonstrate that scSpatialSIM allows users to easily simulate various kernel densities of probability distributions used to create the marked point pattern - points distributed in space with either numeric or categorical features. Using scSpatialSIM, we used four univariate spatial simulation scenarios to compare three different measures for spatial clustering (Ripleys K(r), nearest neighbor G(r), and pair correlation g(r)). We found that Ripleys K(r) identifies the most radii with significant clustering in all four scenarios. Nearest neighbor G(r) only identified all samples as significantly clustered at one radius (r = 0.07) in one simulation scenario (high abundance large cluster size). Pair correlation g(r) was better able to detect significant clustering at low radii when abundance was low. ConclusionsVignettes developed for scSpatialSIM cover the creation of single-type and multi-type spatial single-cell molecular data, as well as how these simulated data can be used with other R packages, such as spatialTIME, to derive spatial statistics. Development of this package is crucial for furthering our understanding of the power of existing methods and the development of novel applications to assess the spatial contexture of tissues by providing an objective platform for simulating spatial single-cell molecular data.

cell biology↗

Molecular subtypes of high grade serous ovarian cancer across racial groups and gene expression platforms

IntroductionHigh-grade serous carcinoma (HGSC) gene expression subtypes are associated with differential survival. We characterized HGSC gene expression in Black individuals and considered whether gene expression differences by race may contribute to poorer HGSC survival among Black versus non-Hispanic White individuals. MethodsWe included newly generated RNA-Seq data from Black and White individuals, and array-based genotyping data from four existing studies of White and Japanese individuals. We assigned subtypes using K-means clustering. Cluster- and dataset-specific gene expression patterns were summarized by moderated t-scores. We compared cluster-specific gene expression patterns across datasets by calculating the correlation between the summarized vectors of moderated t-scores. Following mapping to The Cancer Genome Atlas (TCGA)-derived HGSC subtypes, we used Cox proportional hazards models to estimate subtype-specific survival by dataset. ResultsCluster-specific gene expression was similar across gene expression platforms. Comparing the Black study population to the White and Japanese study populations, the immunoreactive subtype was more common (39% versus 23%-28%) and the differentiated subtype less common (7% versus 22%-31%). Patterns of subtype-specific survival were similar between the Black and White populations with RNA-Seq data; compared to mesenchymal cases, the risk of death was similar for proliferative and differentiated cases and suggestively lower for immunoreactive cases (Black population HR=0.79 [0.55, 1.13], White population HR=0.86 [0.62, 1.19]). ConclusionsA single, platform-agnostic pipeline can be used to assign HGSC gene expression subtypes. While the observed prevalence of HGSC subtypes varied by race, subtype-specific survival was similar. Statement of SignificanceA single pipeline was used to subtype ovarian high-grade serous carcinoma (HGSC) with array-based or RNA-Seq gene expression data. Subtype distributions differed by race, but subtype-specific survival was similar across racial groups.

cancer biology↗