Search bioRxiv⌕ Search

Biology subjects

Janakarajan, N.

Publications and source records attributed to Janakarajan, N..

2 recordsLinked to original sources

Signature Informed Sampling for Transcriptomic Data

With machine learning taking over biomedical applications, working with transcriptomic data on supervised learning tasks is challenging due to high dimensionality, low patient numbers and class imbalances. Machine learning models tend to overfit these data and do not generalise well on out-of-distribution samples. Data augmentation strategies help alleviate this by introducing synthetic data points and acting as regularisers. However, existing approaches are either computationally intensive, require population parametric estimates or generate insufficiently diverse samples. To address these challenges, we introduce two classes of phenotype driven data augmentation approaches - signature-dependent and signature-independent. The signature-dependent methods assume the existence of distinct gene signatures describing some phenotype and are simple, non-parametric, and novel data augmentation methods. The signature-independent methods are a modification of the established Gamma-Poisson and Poisson sampling methods for gene expression data. As case studies, we apply our augmentation methods to transcriptomic data of colorectal and breast cancer. Through discriminative and generative experiments with external validation, we show that our methods improve patient stratification by 5 - 15% over other augmentation methods in different cases. The study additionally provides insights into the limited benefits of over-augmenting data. The code is hosted on GitHub, and includes a link to the augmented datasets.

bioinformatics↗

SurvBoard: Standardised Benchmarking for Multi-omics Cancer Survival Models

Multi-omics data, which include genomic, transcriptomic, epigenetic, and proteomic data, are gaining increasing importance for determining the clinical outcomes of cancer patients. Several recent studies have evaluated various multi-modal integration strategies for cancer survival prediction, highlighting the need for standardizing model performance results. Addressing this issue, we introduce SurvBoard, a benchmark framework that standardizes key experimental design choices. SurvBoard enables comparisons between single-cancer and pan-cancer data models and assesses the benefits of using patient data with missing modalities. We also address common pitfalls in preprocessing and validating multi-omics cancer survival models. We apply SurvBoard to several exemplary use cases, further confirming that statistical models tend to outperform deep learning methods, especially for metrics measuring survival function calibration. Moreover, most models exhibit better performance when trained in a pan-cancer context and can benefit from leveraging samples for which data of some omics modalities are missing. We provide a web service for model evaluation and to make our benchmark results easily accessible and viewable: https://www.survboard.science/. All code is available on GitHub: https://github.com/BoevaLab/survboard/. All benchmark outputs are available on Zenodo: https://zenodo.org/records/11066227. O_TEXTBOXKey MessagesO_LIWe introduce SurvBoard, a comprehensive benchmarking framework for the standardized evaluation of multi-omics cancer survival models. SurvBoard provides an easily accessible platform for the reproducible comparison of models trained on single-cancer and pan-cancer datasets. The platform addresses issues such as the impact of missing modalities and variability in experimental setups. SurvBoard integrates data from four major cancer programs-TCGA, ICGC, TARGET, and METABRIC-to ensure a comprehensive evaluation across diverse types of cancer and research centers. C_LIO_LISurvBoard results confirm that statistical models generally outperform deep learning models in survival function calibration. We also find that pan-cancer training enhances model performance and that models benefit from incorporating data with missing modalities. C_LIO_LISurvBoard includes a web service that allows researchers to submit models for benchmarking and evaluation. A leaderboard is accessible via https://survboard.science/ to promote transparency and the continuous assessment of models performance. C_LI C_TEXTBOX

bioinformatics↗