Search bioRxivSearch

Biology subjects

Martens, L.

Publications and source records attributed to Martens, L..

6 recordsLinked to original sources

Accurate peptide fragmentation predictions allow data driven approaches to replace and improve upon proteomics search engine scoring functions

The use of post-processing tools to maximize the information gained from a proteomics search engine is widely accepted and used by the community, with the most notable example being Percolator - a semi-supervised machine learning model which learns a new scoring function for a given dataset. The usage of such tools is however bound to the search engines scoring scheme, which doesnt always make full use of the intensity information present in a spectrum. By leveraging another machine learning-based tool, MS2PIP, we aim to overcome this obstacle. MS2PIP predicts fragment ion peak intensities. We show how comparing these intensities to annotated experimental spectra by calculating direct similarity metrics rather than the more common peak counting or explained intensities summing provides enough information for a tool such as Percolator to accurately separate two classes of PSMs, recovering more information out of the data while maintaining control of statistics such as the false discovery rate.

bioinformatics

Unbiased dynamic characterization of RNA-protein interactions by OOPS

Current methods for the identification of RNA-protein interactions require a quantity and quality of sample that hinders their application, especially for dynamic biological systems or when sample material is limiting. Here, we present a new approach to enrich RNA-Binding Proteins (RBPs): Orthogonal Organic Phase Separation (OOPS), which is compatible with downstream proteomics and RNA sequencing. OOPS enables recovery of RBPs and free protein, or protein-bound RNA and free RNA, from a single sample in an unbiased manner. By applying OOPS to human cell lines, we extract the majority of known RBPs, and importantly identify additional novel RBPs, including those from previously under-represented cellular compartments. The high yield and unbiased nature of OOPS facilitates its application in both dynamic and inaccessible systems. Thus, we have identified changes in RNA-protein interactions in mammalian cells following nocodazole cell-cycle arrest, and defined the first bacterial RNA-interactome. Overall, OOPS provides an easy-to-use and flexible technique that opens new opportunities to characterize RNA-protein interactions and explore their dynamic behaviour.

cell biology

Comprehensive and empirical evaluation of machine learning algorithms for LC retention time prediction

Liquid chromatography is a core component of almost all mass spectrometric analyses of (bio)molecules. Because of the high-throughput nature of mass spectrometric analyses, the interpretation of these chromatographic data increasingly relies on informatics solutions that attempt to predict an analytes retention time. The key components of such predictive algorithms are the features these are supplies with, and the actual machine learning algorithm used to fit the model parameters.\n\nWe here therefore evaluate the performance of seven machine learning algorithms on 36 distinct metabolomics data sets, using two distinct feature sets. Interestingly, the results show that no single learning algorithm performs optimally for all data sets, with different algorithm types achieving top performance for different types of analytes or different protocols. Our results can thus be used to find an optimal retention time prediction algorithm for specific analytes or protocols. Importantly, however, our results also show that blending different types of models together decreases the error on outliers, indicating that the combination of several approaches holds substantial promise for the development of more generic, high-performing algorithms.

bioinformatics

Beyond the ribosome: proteome-wide secretability studies using SECRiFY

While transcriptome- and proteome-wide technologies to assess processes in protein biogenesis are now widely available, we still lack global approaches to assay post-ribosomal biogenesis events, in particular those occurring in the eukaryotic secretory system. We here developed a method, SECRiFY, to simultaneously assess the secretability of >105 protein fragments by two yeast species, S. cerevisiae and P. pastoris, using custom fragment libraries, surface display and a sequencing-based readout. Screening human proteome fragments with a median size of 50 - 100 amino acids, we generated datasets that enable datamining into protein features underlying secretability, revealing a striking role for intrinsic disorder and chain flexibility. SECRiFY is the first methodology that generates sufficient amounts of annotated data for advanced machine learning methods to deduce secretability predictors. The finding that secretability is indeed a learnable feature of protein sequences is of significant impact in the broad area of recombinant protein expression and de novo protein design.

molecular biology

SQANTI: extensive characterization of long read transcript sequences for quality control in full-length transcriptome identification and quantification

High-throughput sequencing of full-length transcripts using long reads has paved the way for the discovery of thousands of novel transcripts, even in very well annotated organisms as mice and humans. Nonetheless, there is a need for studies and tools that characterize these novel isoforms. Here we present SQANTI, an automated pipeline for the classification of long-read transcripts that computes 47 descriptors that can be used to assess the quality of the data and of the preprocessing pipelines. We applied SQANTI to a neuronal mouse transcriptome using PacBio long reads and illustrate how the tool is effective in readily describing the composition of and characterizing the full-length transcriptome. We perform extensive evaluation of ToFU PacBio transcripts by PCR to reveal that an important number of the novel transcripts are technical artifacts of the sequencing approach, and that SQANTI quality descriptors can be used to engineer a filtering strategy to remove them. Most novel transcripts in this curated transcriptome are novel combinations of existing splice sites, result more frequently in novel ORFs than novel UTRs and are enriched in both general metabolic and neural specific functions. We show that these new transcripts have a major impact in the correct quantification of transcript levels by state-of-the-art short-read based quantification algorithms. By comparing our iso-transcriptome with public proteomics databases we find that alternative isoforms are elusive to proteogenomics detection and are variable in protein changes with respect to the principal isoform of their genes. SQANTI allows the user to maximize the analytical outcome of long read technologies by providing the tools to deliver quality-evaluated and curated full-length transcriptomes. SQANTI is available at https://bitbucket.org/ConesaLab/sqanti.

bioinformatics

Mass spectrometrists should search for all peptides, but assess only the ones they care about

In shotgun proteomics identified mass spectra that are deemed irrelevant to the scientific hypothesis are often discarded. Noble (2015) 1 therefore urged researchers to remove irrelevant peptides from the database prior to searching to improve statistical power. We here however, argue that both the classical as well as Nobles revised method produce suboptimal peptide identifications and have problems in controlling the false discovery rate (FDR). Instead, we show that searching for all expected peptides, and removing irrelevant peptides prior to FDR calculation results in more reliable identifications at controlled FDR level than the classical strategy that discards irrelevant peptides post FDR calculation, or than Nobles strategy that discards irrelevant peptides prior to searching.

bioinformatics