bioRxiv · 10.1101/2024.10.10.617527
Semi-supervised Omics Factor Analysis (SOFA) disentangles known sources of variation from latent factors in multi-omics data
Abstract
A fundamental design pattern in biomolecular studies is to assay the same set of samples (organisms, tissue biopsies, or individual cells) by multiple different omics assays. Group Factor Analysis (GFA) and its adaptation to high-dimensional settings, Multi-Omics Factor Analysis (MOFA), are widely used as a first-line approach to analyse such data and are effective in detecting patterns of correlation, organize them into so-called latent factors, and identify common and assay-specific factors. However, in many applications a subset of the found factors just rediscovers already known covariates (e.g., disease subtypes, environmental covariates) while others may represent genuine novelty. Here, we present Semi-supervised Omics Factor Analysis (SOFA), a method that incorporates known covariates into the model upfront and focuses the factor discovery on novel sources of variation. We show SOFAs effectiveness for discovering novel patterns by applying it to cancer, brain development and heart failure multi-omic data sets.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Capraz, T., Vöhringer, H. S., Huber, W.. 2024-10-14. Semi-supervised Omics Factor Analysis (SOFA) disentangles known sources of variation from latent factors in multi-omics data. https://doi.org/10.1101/2024.10.10.617527
Cite the original work for its findings. Save a collection to share your selection of sources.