Search bioRxivSearch

Biology subjects

Lozoya, O. A.

Publications and source records attributed to Lozoya, O. A..

2 recordsLinked to original sources

Single-cell analyses identify tobacco smoke exposure-associated, dysfunctional CD16+ CD8 T cells with high cytolytic potential in peripheral blood

Tobacco smoke exposure has been found to impact immune response, leukocyte subtypes, DNA methylation, and gene expression in human whole blood. Analysis with single cell technologies will resolve smoking associated (sub)population compositions, gene expression differences, and identification of rare subtypes masked by bulk fraction data. To characterize smoking-related gene expression changes in primary immune cells, we performed single-cell RNA sequencing (scRNAseq) on >45,000 human peripheral blood mononuclear cells (PBMCs) from smokers (n=4) and nonsmokers (n=4). Major cell type population frequencies showed strong correlation between scRNAseq and mass cytometry. Transcriptomes revealed an altered subpopulation of Natural Killer (NK)-like T lymphocytes in smokers, which expressed elevated levels of FCGR3A (gene encoding CD16) compared to other CD8 T cell subpopulations. Relatively rare in nonsmokers (median: 1.8%), the transcriptionally unique subset of CD8 T cells comprised 7.3% of PBMCs in smokers. Mass cytometry confirmed a significant increase (p = 0.03) in the frequency of CD16+ CD8 T cells in smokers. The majority of CD16+ CD8 T cells were CD45RA positive, indicating an effector memory re-expressing CD45RA T cell (TEMRA) phenotype. We expect that cigarette smoke alters CD8 T cell composition by shifting CD8 T cells toward differentiated functional states. Pseudotemporal ordering of CD8 T cell clusters revealed that smokers cells were biased toward later pseudotimes, and characterization of established markers in CD8 T cell subsets indicates a higher frequency of terminally differentiated cells in smokers than in nonsmokers, which corresponded with a lower frequency in naive CD8 T cells. Consistent with an end-stage TEMRA phenotype, FCGR3A-expressing CD8 T cells were inferred as the most differentiated cluster by pseudotime analysis and expressed markers linked to senescence. Examination of differentially expressed genes in other PBMCs uncovered additional senescence-associated genes in CD4 T cells, NKT cells, NK cells, and monocytes. We also observed elevated Tregs, inducers of T cell senescence, in smokers. Taken together, our results suggest smoking-induced, senescence-associated immune cell dysregulation contributes to smoking-mediated pathologies.

immunology

Patterns, Profiles, and Parsimony: dissecting transcriptional signatures from minimal single-cell RNA-seq output with SALSA

Single-cell RNA sequencing (scRNA-seq) technologies have precipitated the development of bioinformatic tools to reconstruct cell lineage specification and differentiation processes with single-cell precision. However, start-up costs and data volumes currently required for statistically reproducible insight remain prohibitively expensive, preventing scRNA-seq technologies from becoming mainstream. Here, we introduce single-cell amalgamation by latent semantic analysis (SALSA), a versatile workflow to address those issues from a data science perspective. SALSA is an integrative and systematic methodology that introduces matrix focusing, a parametric frequentist approach to identify fractions of statistically significant and robust data within single-cell expression matrices. SALSA then transforms the focused matrix into an imputable mix of data-positive and data-missing information, projects it into a latent variable space using generalized linear modelling, and extracts patterns of enrichment. Last, SALSA leverages multivariate analyses, adjusted for rates of library-wise transcript detection and cluster-wise gene representation across latent patterns, to assign individual cells under distinct transcriptional profiles via unsupervised hierarchical clustering. In SALSA, cell type assignment relies exclusively on genes expressed both robustly, relative to sequencing noise, and differentially, among latent patterns, which represent best-candidates for confirmatory validation assays. To benchmark how SALSA performs in experimental settings, we used the publicly available 10X Genomics PBMC 3K dataset, a pre-curated silver standard comprising 2,700 single-cell barcodes from human frozen peripheral blood with transcripts aligned to 16,634 genes. SALSA identified at least 7 distinct transcriptional profiles in PBMC 3K based on <500 differentially expressed Profiler genes determined agnostically, which matched expected frequencies of dominant cell types in peripheral blood. We confirmed that each transcriptional profile inferred by SALSA matched known expression signatures of blood cell types based on surveys of 15 landmark genes and other supplemental markers. SALSA was able to resolve transcriptional profiles from only [~]9% of the total count data accrued, spread across <0.5% of the PBMC 3K expression matrix real estate (16,634 genes x 2,700 cells). In conclusion, SALSA amalgamates scRNA-seq data in favor of reproducible findings. Furthermore, by extracting statistical insight at lower experimental costs and computational workloads than previously reported, SALSA represents an alternative bioinformatics strategy to make single-cell technologies affordable and widespread.

bioinformatics