Search bioRxivSearch

Biology subjects

Chai, H.

Publications and source records attributed to Chai, H..

2 recordsLinked to original sources

Integrating multi-omics data with deep learning for predicting cancer prognosis

BackgroundGenomic information is nowadays widely used for precise cancer treatments. Since the individual type of omics data only represents a single view that suffers from data noise and bias, multiple types of omics data are required for accurate cancer prognosis prediction. However, it is challenging to effectively integrate multi-omics data due to the large number of redundant variables but relatively small sample size. With the recent progress in deep learning techniques, Autoencoder was used to integrate multi-omics data for extracting representative features. Nevertheless, the generated model is fragile from data noises. Additionally, previous studies usually focused on individual cancer types without making comprehensive tests on pan-cancer. Here, we employed the denoising Autoencoder to get a robust representation of the multi-omics data, and then used the learned representative features to estimate patients risks. ResultsBy applying to 15 cancers from The Cancer Genome Atlas (TCGA), our method was shown to improve the C-index values over previous methods by 6.5% on average. Considering the difficulty to obtain multi-omics data in practice, we further used only mRNA data to fit the estimated risks by training XGboost models, and found the models could achieve an average C-index value of 0.627. As a case study, the breast cancer prognosis prediction model was independently tested on three datasets from the Gene Expression Omnibus (GEO), and shown able to significantly separate high-risk patients from low-risk ones (C-index>0.6, p-values<0.05). Based on the risk subgroups divided by our method, we identified nine prognostic markers highly associated with breast cancer, among which seven genes have been proved by literature review. ConclusionOur comprehensive tests indicated that we have constructed an accurate and robust framework to integrate multi-omics data for cancer prognosis prediction. Moreover, it is an effective way to discover cancer prognosis-related genes.

bioinformatics

Imputing missing RNA-seq data from DNA methylation by using transfer learning based-deep neural network

BackgroundGene expression plays a key intermediate role in linking molecular features at DNA level and phenotype. However, due to various limitations in experiments, the RNA-seq data is missing in many samples while there exists high-quality of DNA methylation data. As DNA methylation is an important epigenetic modification to regulate gene expression, it can be used to predict RNA-seq data. For this purpose, many methods have been developed. A common limitation of these methods is that they mainly focus on single cancer dataset, and do not fully utilize information from large pan-cancer dataset. ResultsHere, we have developed a novel method to impute missing gene expression data from DNA methylation data through transfer learning-based neural network, namely TDimpute. In the method, the pan-cancer dataset from The Cancer Genome Atlas (TCGA) was utilized for training a general model, which was then fine-tuned on the specific cancer dataset. By testing on 16 cancer datasets, we found that our method significantly outperforms other state-of-the-art methods in imputation accuracy with 7%-11% increase under different missing rates. The imputed gene expression was further proved to be useful for downstream analyses, including the identification of both methylation-driving and prognosis-related genes, clustering analysis, and survival analysis on the TCGA dataset. More importantly, our method was indicated to be useful for general purpose by the independent test on the Wilms tumor dataset from the Therapeutically Applicable Research To Generate Effective Treatments (TARGET) project. ConclusionsTDimpute is an effective method for RNA-seq imputation with limited training samples.

bioinformatics