Search bioRxivSearch

Biology subjects

Degroeve, S.

Publications and source records attributed to Degroeve, S..

3 recordsLinked to original sources

Accurate peptide fragmentation predictions allow data driven approaches to replace and improve upon proteomics search engine scoring functions

The use of post-processing tools to maximize the information gained from a proteomics search engine is widely accepted and used by the community, with the most notable example being Percolator - a semi-supervised machine learning model which learns a new scoring function for a given dataset. The usage of such tools is however bound to the search engines scoring scheme, which doesnt always make full use of the intensity information present in a spectrum. By leveraging another machine learning-based tool, MS2PIP, we aim to overcome this obstacle. MS2PIP predicts fragment ion peak intensities. We show how comparing these intensities to annotated experimental spectra by calculating direct similarity metrics rather than the more common peak counting or explained intensities summing provides enough information for a tool such as Percolator to accurately separate two classes of PSMs, recovering more information out of the data while maintaining control of statistics such as the false discovery rate.

bioinformatics

Unbiased dynamic characterization of RNA-protein interactions by OOPS

Current methods for the identification of RNA-protein interactions require a quantity and quality of sample that hinders their application, especially for dynamic biological systems or when sample material is limiting. Here, we present a new approach to enrich RNA-Binding Proteins (RBPs): Orthogonal Organic Phase Separation (OOPS), which is compatible with downstream proteomics and RNA sequencing. OOPS enables recovery of RBPs and free protein, or protein-bound RNA and free RNA, from a single sample in an unbiased manner. By applying OOPS to human cell lines, we extract the majority of known RBPs, and importantly identify additional novel RBPs, including those from previously under-represented cellular compartments. The high yield and unbiased nature of OOPS facilitates its application in both dynamic and inaccessible systems. Thus, we have identified changes in RNA-protein interactions in mammalian cells following nocodazole cell-cycle arrest, and defined the first bacterial RNA-interactome. Overall, OOPS provides an easy-to-use and flexible technique that opens new opportunities to characterize RNA-protein interactions and explore their dynamic behaviour.

cell biology

Comprehensive and empirical evaluation of machine learning algorithms for LC retention time prediction

Liquid chromatography is a core component of almost all mass spectrometric analyses of (bio)molecules. Because of the high-throughput nature of mass spectrometric analyses, the interpretation of these chromatographic data increasingly relies on informatics solutions that attempt to predict an analytes retention time. The key components of such predictive algorithms are the features these are supplies with, and the actual machine learning algorithm used to fit the model parameters.\n\nWe here therefore evaluate the performance of seven machine learning algorithms on 36 distinct metabolomics data sets, using two distinct feature sets. Interestingly, the results show that no single learning algorithm performs optimally for all data sets, with different algorithm types achieving top performance for different types of analytes or different protocols. Our results can thus be used to find an optimal retention time prediction algorithm for specific analytes or protocols. Importantly, however, our results also show that blending different types of models together decreases the error on outliers, indicating that the combination of several approaches holds substantial promise for the development of more generic, high-performing algorithms.

bioinformatics