Search bioRxivSearch

Biology subjects

Käll, L.

Publications and source records attributed to Käll, L..

3 recordsLinked to original sources

Quantitative behavior of protein complexes in human cells

Translational and post-translational control mechanisms in the cell result in widely observable differences between measured gene transcription and protein abundances. Herein, protein complexes are among the most tightly controlled entities by selective degradation of their individual proteins. They furthermore act as control hubs that regulate highly important processes in the cell and exhibit a high functional diversity due to their ability to change their composition and their structure. To better understand and predict these functional states, extensive characterization of complex composition, behavior, and abundance is necessary. Mass spectrometry provides an unbiased approach to directly determine protein abundances across cell populations and thus to profile a comprehensive abundance map of proteins. We investigated the behavior of protein subunits in known complexes by comparing their abundance profiles across up to 140 cell types available in ProteomicsDB. After thorough assessment of different randomization methods and statistical scoring algorithms, we developed a computational tool to quantify the significance of concurrent profiles within a complex, therefore providing insights into the conservation of their composition across human cell types. We identified the intrinsic structures in complex behavior that allow to determine which proteins orchestrate complex function. This analysis can be extended to investigate common profiles within arbitrary protein groups. With the CoExpresso web service, we offer a potent scoring scheme to assess proteins for their co-regulation and thereby offer insight into their potential for forming functional groups like protein complexes. CoExpresso can be accessed through http://computproteomics/Apps/CoExpresso. Source code and R scripts for database generation are available at https://bitbucket.org/veitveit/coexpresso.\n\nAuthor summaryMany proteins form multi-functional assemblies called protein complexes instead of working as singly units. These complexes control most processes in the cell making the full characterization of their behavior inevitable to understand cellular control mechanisms. Detailed knowledge about complex behavior will elucidate biomarkers and drug targets that exhibit and correct aberrant cell states, respectively. We investigated abundance changes of the protein complex components over more than 100 different human cell types. By using statistical scoring models, we estimated the evidence for the co-regulation of the proteins and revealed which proteins form subunits with impact on complex function and composition. By providing the interactive web service CoExpresso, any combination of proteins can be tested for their co-regulation in human cells.

bioinformatics

Integrated identification and quantification error probabilities for shotgun proteomics

Protein quantification by label-free shotgun proteomics experiments is plagued by a multitude of error sources. Typical pipelines for identifying differentially expressed proteins use intermediate filters in an attempt to control the error rate. However, they often ignore certain error sources and, moreover, regard filtered lists as completely correct in subsequent steps. These two indiscretions can easily lead to a loss of control of the false discovery rate (FDR). We propose a probabilistic graphical model, Triqler, that propagates error information through all steps, employing distributions in favor of point estimates, most notably for missing value imputation. The model outputs posterior probabilities for fold changes between treatment groups, highlighting uncertainty rather than hiding it. We analyzed 3 engineered datasets and achieved FDR control and high sensitivity, even for truly absent proteins. In a bladder cancer clinical dataset we discovered 35 proteins at 5% FDR, whereas the original study discovered 1 and MaxQuant/Perseus 4 proteins at this threshold. Compellingly, these 35 proteins showed enrichment for functional annotation terms, whereas the top ranked proteins reported by MaxQuant/Perseus showed no enrichment. The model executes in minutes and is freely available at https://pypi.org/project/triqler/.

bioinformatics

A protein standard that emulates homology for the characterization of protein inference algorithms

A natural way to benchmark the performance of an analytical experimental setup is to use samples of known content, and see to what degree one can correctly infer the content of such a sample from the data. For shotgun proteomics, one of the inherent problems of interpreting data is that the measured analytes are peptides and not the actual proteins themselves. As some proteins share proteolytic peptides, there might be more than one possible causative set of proteins resulting in a given set of peptides and there is a need for mechanisms that infer proteins from lists of detected peptides. A weakness of commercially available samples of known content is that they consist of proteins that are deliberately selected for producing tryptic peptides that are unique to a single protein. Unfortunately, such samples do not expose any complications in protein inference. For a realistic benchmark of protein inference procedures, there is, therefore, a need for samples of known content where the present proteins share peptides with known absent proteins. Here, we present such a standard, that is based on E. coli expressed human protein fragments. To illustrate the usage of this standard, we benchmark a set of different protein inference procedures on the data. We observe that inference procedures excluding shared peptides provide more accurate estimates of errors compared to methods that include information from shared peptides, while still giving a reasonable performance in terms of the number of identified proteins. We also demonstrate that using a sample of known protein content without proteins with shared tryptic peptides can give a false sense of accuracy for many protein inference methods.

bioinformatics