Search bioRxiv⌕ Search

Biology subjects

Roettger, R.

Publications and source records attributed to Roettger, R..

2 recordsLinked to original sources

Variance Analysis of LC-MS Experimental Factors and Their Impact on Machine Learning

BackgroundMachine learning (ML) technologies, especially deep learning (DL), have gained increasing attention in predictive mass spectrometry (MS) for enhancing the data processing pipeline from raw data analysis to end-user predictions and re-scoring. ML models need large-scale datasets for training and re-purposing, which can be obtained from a range of public data repositories. However, applying ML to public MS datasets on larger scales is challenging, as they vary widely in terms of data acquisition methods, biological systems, and experimental designs. ResultsWe aim to facilitate ML efforts in MS data by conducting a systematic analysis of the potential sources of variance in public MS repositories. We also examine how these factors affect ML performance and perform a comprehensive transfer learning to evaluate the benefits of current best practice methods in the field for transfer learning. ConclusionsOur findings show significantly higher levels of homogeneity within a project than between projects, which indicates that its important to construct datasets most closely resembling future test cases, as transferability is severely limited for unseen datasets. We also found that transfer learning, although it did increase model performance, did not increase model performance compared to a non-pre-trained model.

bioinformatics↗

MoSBi: Automated signature mining for molecular stratification and subtyping

The improving access to increasing amounts of biomedical data provides completely new chances for advanced patient stratification and disease subtyping strategies. This requires computational tools that produce uniformly robust results across highly heterogeneous molecular data. Unsupervised machine learning methodologies are able to discover de-novo patterns in such data. Biclustering is especially suited by simultaneously identifying sample groups and corresponding feature sets across heterogeneous omics data. The performance of available biclustering algorithms heavily depends on individual parameterization and varies with their application. Here, we developed MoSBi (Molecular Signature identification using Biclustering), an automated multi-algorithm ensemble approach that integrates results utilizing an error model-supported similarity network. We evaluated the performance of MoSBi on transcriptomics, proteomics and metabolomics data, as well as synthetic datasets covering various data properties. Profiting from multi-algorithm integration, MoSBi identified robust group and disease specific signatures across all scenarios overcoming single algorithm specificities. Furthermore, we developed a scalable network-based visualization of bicluster communities that support biological hypothesis generation. MoSBi is available as an R package and web-service to make automated biclustering analysis accessible for application in molecular sample stratification.

bioinformatics↗