bioRxiv · 10.1101/2021.03.27.437353
Machine learning algorithms in big data analyses identify determinants of insulin gene transcription
Abstract
Machine learning (ML)-workflows enable unprejudiced/robust evaluation of complex datasets. Here, we analyzed over 490,000,000 data points to compare 10 different ML-workflows in a large (N=11,652) training dataset of human pancreatic single-cell (sc-)transcriptomes to identify genes associated with the presence or absence of insulin transcript(s). Prediction accuracy/sensitivity of each ML-workflow were tested in a separate validation dataset (N=2,913). Ensemble ML-workflows, in particular Random Forest ML-algorithm delivered high predictive power (AUC=0.83) and sensitivity (0.98), compared to other algorithms. The transcripts identified through these analyses also demonstrated significant correlation with insulin in bulk RNA-seq data from human islets. The top-10 features, (including IAPP, ADCYAP1, LDHA and SST) common to the three Ensemble ML-workflows were significantly dysregulated in scRNA-seq datasets from Ire-1{beta}-/- mice that demonstrate dedifferentiation of pancreatic {beta}-cells in a model of type 1 diabetes (T1D) and in pancreatic single cells from individuals with type 2 Diabetes (T2D). Our findings provide direct comparison of ML-workflows in big data analyses, identify key determinants of insulin transcription and provide workflows for future analyses.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Wong, W. K., Thorat, V., Joglekar, M. V., Dong, C. X., Lee, H., Bhave, A., Engin, F., Pant, A., Dalgaard, L. T., Bapat, S., Hardikar, A. A.. 2021-03-29. Machine learning algorithms in big data analyses identify determinants of insulin gene transcription. https://doi.org/10.1101/2021.03.27.437353
Cite the original work for its findings. Save a collection to share your selection of sources.