Search bioRxiv⌕ Search

Biology subjects

Ma, S.-L.

Publications and source records attributed to Ma, S.-L..

3 recordsLinked to original sources

In-silico cell sorting revealed granulocyte-specific single-cell-type gene expression from peripheral blood bulk expression data and its application as host response biomarkers to discriminate bacterial and viral infections

Peripheral Blood transcriptome analysis evaluated the bulk transcript abundance (TA) covering all leukocyte cell populations. However, there are 2 main problems in using bulk expression as biomarkers: (1) A long list of differential expression genes (DEGs) was found, and (2) DEGs cannot be attributed to a host response of any specific cell-type. TA assays after conventional cell sorting, as the gold-standard method, is too tedious for routine use. Recently, we showed that by using a ratio-based biomarker, RBB (ratio of two stringently selected genes), it is feasible to interrogate the gene expression of a single cell-type (monocyte and B lymphocyte) in peripheral whole blood (WB) directly. Here, we apply this in-silico cell sorting algorithm (DIRECT LS-TA, Direct Leukocyte Single cell-type Transcript Abundance) to granulocytes in WB samples to reveal RBBs specific to granulocytes. This DIRECT LS-TA approach without the need for cell-sorting was applied to public datasets to differentiate the 2 types of infection (bacterial vs viral infection). The following RBBs measured in WB correlate with the expression of target (numerator) genes in purified granulocytes, thus cell-sorting can be avoided by using these RBBs: ARG1/SRGN, ANXA3/SRGN, RSAD2/SRGN. Together with monocyte DIRECT LS-TA biomarkers, IFI27/PSAP, direct quantification of 4 genes provided optimal differentiation of viral from bacterial infection. Meta-analysis and unsupervised machine learning classification confirmed the superior performance of DIRECT LS-TA biomarkers. These RBBs found by prior In-silico cell-sorting identified pairs of genes that are used to formulate as ratio-based biomarkers (RBBs) to represent gene expression of granulocytes inside whole blood cell-mixture samples which was useful to triage febrile patients into two major categories of febrile diseases between viral and bacterial infection with high degree of sensitivity and specificity.

immunology↗

A map of Non-translated RNA (nt-RNA) junctions in cancer genomes: a database resource of unproductive splicing

BackgroundNon-translated transcripts (nt-RNAs) with frame-shifts or premature termination codons resulting from alternative splicing events (ASE), have been recently found at unexpectedly abundant in transcriptomes of cancer tissue. However, their full genomic spectrum has not yet been fully elucidated. This study comprehensively characterised the expression of signature junctions of these nt-RNA (termed "toxic junctions" here) of both known and novel nt-RNA across multiple cancer types and investigated their potential as biomarkers. MethodsRNA-seq data of [~]6,000 samples, including the tumor and normal samples for 13 cancer types were retrieved from The Cancer Genome Atlas database (TCGA) together with data from Cancer Cell Line Encyclopedia (CCLE) project. Due to the difficulty in quantifying the entire transcript isoform of nt-RNA, we pioneered an algorithm to focus exclusively on the expression of junctional reads, which also circumvented the limitation of non-directional RNA- seq of TCGA data. We showed that the majority of nt-RNA is associated with at least one toxic junction. We built a comprehensive catalogue of known nt-RNA toxic junctions from genome databases. And novel toxic junctions were also identified by a new junction-focused algorithm from the higher quality discovery subsets of TCGA data. Splicing in Ratio (SiR) was used to quantify ASE leading to nt-RNA, enabling: Differential expression analysis between cancer and normal tissue and across cancer types. Identification of different profiles of nt-RNA abundance and various factor which may be the causes of differential nt-RNA abundance and SiR results Identification of specific nt-RNA and toxic junctions that were expressed in various cancer (and/or normal tissue) types. Assessment of nt-RNA and their toxic junction expression as biomarkers or prognosis indicators. ResultsWe profiled the expressed known nt-RNA (toxic) junctions of known transcripts and discovered [~]22,000 novel toxic junctions out of [~]250,000 novel junctions found in the transcriptome data. The expression of nt-RNA was as high as 10% of all transcripts of the corresponding gene in cancer transcriptomes. Interestingly, some signature toxic junctions of nt-RNA are expressed in even higher quantities, e.g. up to 50% or more, which is reminiscent of a heterozygous mutation. We identified distinct patterns between cancer and normal samples, including example of nt-RNA expressing toxic junctions exclusively in normal or tumor samples. Clinically relevant examples included ANXA6 in breast cancer, where the nt-RNA isoform showed significantly higher expression in tumors (p=1.8e-15). In kidney renal clear cell carcinoma (KIRC), a significant isoform switch of ESYT2 based on the RNA-seq data was confirmed. The Kaplan-Meier survival curves showed that samples with the higher expression ratio of ESYT2-L are associated with better survival (p=2.0e-06). Unsupervised clustering showed that SiR results of 150 toxic signatures defined 4 subgroups of patients with different prognosis. Through principal component analysis (PCA), PC1 and PC2 can be used as an independent prognosis biomarkers. nt-RNA accounting for these PCs included splicing factors SRSF3 and CLK1, where CLK1 phosphorylates SRSF3 to promote exon 4 inclusion in both genes. ConclusionsIn summary, the expression profiles of all known and novel toxic junctions were explored using pan-cancer RNA-seq data. A dual 10% rule emerged from this study: [~]10% of novel junctions were toxic junctions associated with nt-RNA, and up to 10% of RNA transcripts inside a cell were also nt-RNA. The SiR metric enables accurate quantification of unproductive splicing and identification of cancer biomarkers. Our findings reveal that unproductive splicing represents functionally important post-transcriptional regulation in cancer. These expression profiles allow researchers to study the expression of nt-RNA signature junctions or novel signature junctions in or near the genes they are interested in, which could provide a new direction for their research. The SRSF3-CLK1 regulatory mechanism provides insights into splicing dysregulation. Our comprehensive toxic junction catalogue serves as a valuable resource, suggesting that targeting unproductive splicing pathways may offer novel therapeutic strategies for cancer treatment. Data availabilityThe catalogue is available on GitHub and UCSC browser. https://github.com/danhuang0909/nt_database for GitHub overview https://genome.ucsc.edu/s/dandan_0909/hg38_all_new_nr for genome browsing of all novel (unannotated) toxic junctions https://genome.ucsc.edu/s/dandan_0909/hg38_5_26 for toxic junctions in known (annotated) nt-RNA.

cancer biology↗

Monocyte single cell-type gene expression measured in peripheral blood by DIRECT LS-TA method: the ratio-based biomarkers of (IFI27/PSAP) showed superior performance than interferon score in triage patients with viral infection

A rapid method to triage febrile patients into different categories of etiologies remains a significant challenge even nowadays, when many molecular tests for pathogens are available. Routine serum protein tests like C-reactive protein and procalcitonin have limited specificity. Host response gene signatures are promising biomarkers but they usually require assaying many genes, e.g. 7 genes are commonly used to calculate the interferon (IFN) score. However, these gene panels fail to capture cell-type-specific host responses. Measuring gene expression of a specified single cell population, like monocytes, offers enhanced biological insight. However, it currently requires laborious cell sorting or costly single-cell sequencing techniques, limiting its clinical applicability. This study aims to develop a simple ratio-based biomarker (RBB) representing monocyte-specific host response to viral infection called DIRECT LS-TA method. A simple ratio of 2 genes (both are shortlist monocyte informative genes) quantified in peripheral blood (PB) samples correlated with gene expression in purified monocytes in the corresponding individual. These RBBs cover 3 interferon-stimulated genes (ISGs): IFI27/PSAP, IFI44L/PSAP and SIGLEC1/PSAP. They are compared to the conventional multi-gene IFN score in the differentiation of viral infection. Public gene expression datasets from NCBI GEO were used to shortlist monocyte-informative genes that can be used as the RBB in PB. The DIRECT LS-TA RBB was calculated as the ratio of the target ISG transcript abundance (TA) to that of another reference gene (PSAP or CTSS) directly quantified from bulk PB data (e.g., Log(IFI27/PSAP) in WB). The correlation (expressed by coefficient of determination, R{superscript 2}) between these DIRECT LS-TA RBBs and the gold-standard target gene TA measured in purified monocytes was assessed. The diagnostic performance of selected RBBs (IFI27/PSAP, IFI44L/PSAP, SIGLEC1/PSAP) was compared against the conventional 8-gene IFN score for differentiating viral infections from controls. Direct LS-TA RBBs measured in PB showed strong correlation with gold-standard gene expression measured in purified monocytes (R2 ranged from 0.53 for the target gene IFI27 to >0.9 for the target gene IFI44L). This high level of correlation supports that this simple RBB (DIRECT LS-TA) method can replace the tedious cell sorting approach to obtain single-cell-type gene expression data. All DIRECT LS-TA results of ISGs were raised during viral infection. The best clinical performance in triaging viral infection patients was achieved by IFI27/PSAP or IFI27/CTSS across all datasets. For example, in the GSE111368 dataset, IFI27/PSAP achieved an AUC of 0.94 (95% CI 0.90-0.97) with 88% sensitivity and 95% specificity, surpassing the IFN scores AUC of 0.90 (95% CI 0.85-0.94) with 79% sensitivity and 93% specificity. ConclusionThe DIRECT LS-TA method, utilizing the format of simple two-gene ratio-based biomarkers like IFI27/PSAP, provides a robust and accurate measure of monocyte-specific interferon pathway activation directly from peripheral blood. The superior performance of the DIRECT LS-TA method makes it a promising, readily implementable tool for clinical triage. Its ability to provide single-cell-type specific information, rapid turnaround using standard qPCR/dPCR technology, and enhanced biological specificity make it a valuable molecular host response assessment.

immunology↗