Search bioRxiv⌕ Search

Biology subjects

Xu, X. S.

Publications and source records attributed to Xu, X. S..

2 recordsLinked to original sources

Multitask Learning of Longitudinal Circulating Biomarkers and Clinical Outcomes: Identification of Optimal Machine-Learning and Deep-Learning Models

Many circulating biomarkers are assessed at different time intervals during clinical studies. Despite of the success of standard joint models in predicting clinical outcomes using low-dimensional longitudinal data (1-2 biomarkers), significant computational challenges are encountered when applying these techniques to high-dimensional biomarker datasets. Modern machine- or deep-learning models show potential for multiple biomarker processes, but systematic evaluations and applications to high-dimensional data in the clinical settings have yet to be reported. We aimed to enhance the scalability of joint modeling and provide guidance on optimal approaches for high-dimensional biomarker data and outcomes. We evaluated multiple deep-learning and machine-learning models using 24 clinical biomarkers and survival data from the SQUIRE trial, a phase 3 randomized clinical trial investigating necitumumab and standard gemcitabine/cisplatin treatment in patients with squamous non-small-cell lung cancer (NSCLC). Overall, we confirmed that longitudinal models enabled more accurate prediction of patients survival compared to those solely based on baseline information. Coupling multivariate functional principal component analysis (MFPCA) with Cox regression (MFPCA-Cox) provided the highest predictive discrimination and accuracy for the NSCLC patients with AUC values of 0.7 - >0.8 at various landmark time points and prediction timeframes, outperforming recent advanced Transformer and convolutional neural network deep-learning algorithms (TransformerJM and Match-Net, respectively). In conclusion, we identified that MFPCA-Cox represents a robust and versatile joint modeling algorithm for high-dimensional biomarker longitudinal data with irregular and missing data, capturing complex relationships within the data, yielding accurate predictions for both longitudinal biomarkers and survival outcomes, and gaining insights into the underlying dynamics.

pharmacology and toxicology↗

Mutually exclusive spectral biclustering and its applications in cancer subtyping

Many soft biclustering algorithms have been developed and applied to various biological and biomedical data analyses. However, until now, few mutually exclusive (hard) biclustering algorithms have been proposed although they can be extremely useful for identify disease or molecular subtypes based on genomic or transcriptomic data. We considered the biclustering problem of expression matrices as a bipartite graph partitioning problem and developed a novel biclustering algorithm, MESBC, based on Dhillons spectral method to detect mutually exclusive biclusters. MESBC simultaneously detects relevant features (genes) and corresponding subgroups, and therefore automatically uses the signature features for each subtype to perform the clustering, improving the clustering performance. MESBC could accurately detect the pre-specified biclusters in simulations, and the identified biclusters were highly consistent with the true labels. Particularly, in setting with high noise, MESBC outperformed existing NMF and Dhillons method and provided markedly better accuracy. Analysis of two TCGA datasets (LUAD and BRAC cohorts) revealed that MESBC provided similar or more accurate prognostication (i.e., smaller p value) for overall survival in patients with breast and lung cancer, respectively, compared to the existing, gold-standard subtypes for breast (PAM50) and lung cancer (integrative clustering). In the TCGA lung cancer patients, MESBC detected two clinically relevant, rare subtypes that other biclustering or integrative clustering algorithms could not detect. These findings validated our hypothesis that MESBC could improve molecular subtyping in cancer patients and potentially facilitate better individual patient management, risk stratification, patient selection, therapeutic assignments, as well as better understanding gene signatures and molecular pathways for development of novel therapeutic agents.

bioinformatics↗