Search bioRxiv⌕ Search

Biology subjects

Ji, B.-Y.

Publications and source records attributed to Ji, B.-Y..

3 recordsLinked to original sources

Multi-view graph learning for deciphering the dominant cell communication assembly of downstream functional events from single-cell RNA-seq data

Cell-cell communications (CCCs) from multiple sender cells collaboratively affect downstream functional events in receiver cells, thus influencing cell phenotype and function. How to rank the importance of these CCCs and find the dominant ones in a specific downstream functional event has great significance for deciphering various physiological and pathogenic processes. To date, several computational methods have been developed to focus on the identification of cell types that communicate with enriched ligand-receptor interactions from single-cell RNA-seq (scRNA-seq) data, but to the best of our knowledge, all of them lack the ability to identify the communicating cell type pairs that play a major role in a specific downstream functional event, which we call it "dominant cell communication assembly (DCA)". Here, we proposed scDCA, a multi-view graph learning method for deciphering DCA from scRNA-seq data. scDCA is based on a multi-view CCC network by constructing different cell type combinations at single-cell resolution. Multi-view graph convolution network was further employed to reconstruct the expression pattern of target genes or the functional states of receiver cells. The DCA was subsequently identified by interpreting the model with the attention mechanism. scDCA was verified in a real scRNA-seq cohort of advanced renal cell carcinoma, accurately deciphering the DCA that affect the expression patterns of the critical immune genes and functional states of malignant cells. Furthermore, scDCA also accurately explored the alteration in cell communication under clinical intervention by comparing the DCA for certain cytotoxic factors between patients with and without immunotherapy. scDCA is free available at: https://github.com/pengsl-lab/scDCA.git.

bioinformatics↗

SpaCCC: Large language model-based cell-cell communication inference for spatially resolved transcriptomic data

Drawing parallels between linguistic constructs and cellular biology, large language models (LLMs) have achieved remarkable success in diverse downstream applications for single-cell data analysis. However, to date, it still lacks methods to take advantage of LLMs to infer ligand-receptor (LR)-mediated cell-cell communications for spatially resolved transcriptomic data. Here, we propose SpaCCC to facilitate the inference of spatially resolved cell-cell communications, which relies on our fine-tuned single-cell LLM and functional gene interaction network to embed ligand and receptor genes expressed in interacting individual cells into a unified latent space. The LR pairs with a significant closer distance in latent space are taken to be more likely to interact with each other. After that, the molecular diffusion and permutation test strategies are respectively employed to calculate the communication strength and filter out communications with low specificities. The benchmarked performance of SpaCCC is evaluated on real single-cell spatial transcriptomic datasets with remarkable superiority over other methods. SpaCCC also infers known LR pairs concealed by existing aggregative methods and then identifies communication patterns for specific cell types and their signalling pathways. Furthermore, spaCCC provides various cell-cell communication visualization results at both single-cell and cell type resolution. In summary, spaCCC provides a sophisticated and practical tool allowing researchers to decipher spatially resolved cell-cell communications and related communication patterns and signalling pathways based on spatial transcriptome data.

bioinformatics↗

HyperVR--A hybrid prediction framework for virulence factors and antibiotic resistance genes in microbial data

Infectious diseases, particularly bacterial infections, are emerging at an unprecedented rate, posing a serious challenge to public health and the global economy. Different virulence factors (VFs) work in concert to enable pathogenic bacteria to successfully adhere, reproduce and cause damage to host cells, and antibiotic resistance genes (ARGs) allow pathogens to evade otherwise curable treatments. To understand the causal relationship between microbiome composition, function and disease, both VFs and ARGs in microbial data must be identified. Most existing computational models cannot simultaneously identify VFs or ARGs, hindering the related research. The best hit approaches are currently the main tools to identify VFs and ARGs concurrently; yet they usually have high false-negative rates and are very sensitive to the cut-off thresholds. In this work, we proposed a hybrid computational framework called HyperVR to predict VFs and ARGs at the same time. Specifically, HyperVR integrates key genetic features and then stacks classical ensemble learning methods and deep learning for training and prediction. HyperVR accurately predicts VFs, ARGs and negative genes (neither VFs nor ARGs) simultaneously, with both high precision (>0.91) and recall (>0.91) rates. Also, HyperVR keeps the flexibility to predict VFs or ARGs individually. Regarding novel VFs and ARGs, the VFs and ARGs in metagenomic data, and pseudo VFs and ARGs (gene fragments), HyperVR has shown good prediction, outperforming the current state-of-the-art predition tools and best hit approaches in terms of precision and recall. HyperVR is a powerful tool for predicting VFs and ARGs simultaneously by using only gene sequences and without strict cut-off thresholds, hence making prediction straightforward and accurate.

bioinformatics↗