bioRxiv · 10.1101/2024.10.31.621233
Exploring the latent space of transcriptomic data withtopic modeling
Abstract
The availability of high-dimensional transcriptomic datasets is increasing at a tremendous pace, together with the need for suitable computational tools. Clustering and dimensionality reduction methods are popular go-to methods to identify basic structures in these datasets. At the same time, different topic modeling techniques have been developed to organize the deluge of available data of natural language using their latent topical structure. This paper leverages the statistical analogies between text and transcriptomic datasets to compare different topic modeling methods when applied to gene expression data. Specifically, we test their accuracy in the specific task of discovering and reconstructing the tissue structure of the human transcriptome and distinguishing healthy from cancerous tissues. We examine the properties of the latent space recovered by different methods, highlight their differences, and the pros and cons of the methods across different tasks. Finally, we show that the latent topic space can be a useful embedding space, where a basic neural network classifier can annotate transcriptomic profiles with high accuracy.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Valle, F., Caselle, M., Osella, M.. 2024-11-03. Exploring the latent space of transcriptomic data withtopic modeling. https://doi.org/10.1101/2024.10.31.621233
Cite the original work for its findings. Save a collection to share your selection of sources.