bioRxiv · 10.1101/2023.05.28.542669
Transformer with Convolution and Graph-Node co-embedding: An accurate and interpretable vision backbone for predicting gene expressions from local histopathological image
Abstract
Inferring gene expressions from histopathological images has always been a fascinating but challenging task due to the huge differences between the two modal data. Previous works have used modified DenseNet121 to encode the local images and make gene expression predictions. And later works improved the prediction accuracy of gene expression by incorporating the coordinate information from images and using all spots in the tissue region as input. While these methods were limited in use due to model complexity, large demand on GPU memory, and insufficient encoding of local images, thus the results had low interpretability, relatively low accuracy, and over-smooth prediction of gene expression among neighbor spots. In this paper, we propose TCGN, (Transformer with Convolution and Graph-Node co-embedding method) for gene expression prediction from H&E stained pathological slide images. TCGN consists of convolutional layers, transformer encoders, and graph neural networks, and is the first to integrate these blocks in a general and interpretable computer vision backbone for histopathological image analysis. We trained TCGN and compared its performance with three existing methods on a publicly available spatial transcriptomic dataset. Even in the absence of the coordinates information and neighbor spots, TCGN still outperformed the existing methods by 5% and achieved 10 times higher prediction accuracy than the counterpart model. Besides its higher accuracy, our model is also small enough to be run on a personal computer and does not need complex building graph preprocessing compared to the existing methods. Moreover, TCGN is interpretable in recognizing special cell morphology and cell-cell interactions compared to models using all spots as input that are not interpretable. A more accurate omics information prediction from pathological images not only links genotypes to phenotypes so that we can predict more biomarkers that are expensive to test from histopathological images that are low-cost to obtain, but also provides a theoretical basis for future modeling of multi-modal data. Our results support that TCGN is a useful tool for inferring gene expressions from histopathological images and other potential histopathological image analysis studies. HighlightsO_LIFirst deep learning model to integrate CNN, GNN, and transformer for image analysis C_LIO_LIAn interpretable model that uses cell morphology and organizations to predict genes C_LIO_LIHigher gene expression prediction accuracy without global information C_LIO_LIAccurately predicted genes are related to immune escape and abnormal metabolism C_LIO_LIPredict important biomarkers for breast cancer accurately from cheaper images C_LI Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=123 SRC="FIGDIR/small/542669v1_ufig1.gif" ALT="Figure 1"> View larger version (51K): org.highwire.dtl.DTLVardef@1c6a4dforg.highwire.dtl.DTLVardef@7266d4org.highwire.dtl.DTLVardef@bcfd0aorg.highwire.dtl.DTLVardef@188d452_HPS_FORMAT_FIGEXP M_FIG C_FIG
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xiao, X., Kong, Y., Wang, Z., Lu, H.. 2023-05-30. Transformer with Convolution and Graph-Node co-embedding: An accurate and interpretable vision backbone for predicting gene expressions from local histopathological image. https://doi.org/10.1101/2023.05.28.542669
Cite the original work for its findings. Save a collection to share your selection of sources.