Search bioRxivSearch

Biology subjects

Lilley, K. S.

Publications and source records attributed to Lilley, K. S..

6 recordsLinked to original sources

LOPIT-DC: A simpler approach to high-resolution spatial proteomics

Hyperplexed Localisation of Organelle Proteins by Isotope Tagging (hyperLOPIT) is a well-established method for studying protein subcellular localisation in complex biological samples. As a simpler alternative we developed a second workflow named Localisation of Organelle Proteins by Isotope Tagging after Differential ultraCentrifugation (LOPIT-DC) which is faster and less resource-intensive. We present the most comprehensive high-resolution mass spectrometry-based human dataset to date and deliver a flexible set of subcellular proteomics protocols for sample preparation and data analysis. For the first time, we methodically compare these two different mass spectrometry-based spatial proteomics methods within the same study and also apply QSep, the first tool that objectively and robustly quantifies subcellular resolution in spatial proteomics data. Using both approaches we highlight suborganellar resolution and isoform-specific subcellular niches as well as the locations of large protein complexes and proteins involved in signalling pathways which play important roles in cancer and metabolism. Finally, we showcase an extensive analysis of the multilocalising proteome identified via both methods.

biochemistry

Assessing sub-cellular resolution in spatial proteomics experiments

The sub-cellular localisation of a protein is vital in defining its function, and a proteins mis-localisation is known to lead to adverse effect. As a result, numerous experimental techniques and datasets have been published, with the aim of deciphering the localisation of proteins at various scales and resolutions, including high profile mass spectrometry-based efforts. Here, we present a meta-analysis assessing and comparing the sub-cellular resolution of 29 such mass spectrometry-based spatial proteomics experiments using a newly developed tool termed QSep. Our goal is to provide a simple quantitative report of how well spatial proteomics resolve the sub-cellular niches they describe to inform and guide developers and users of such methods.

bioinformatics

Biotinylation by proximity labelling favours unfolded proteins

Folded enzymes are essential for life, but there is limited in vivo information about how locally unfolded protein regions contribute to biological functions. Intrinsically Disordered Regions (IDRs) are enriched in disease-linked and multiply post-translationally modified proteins. The extent of foldability of predicted IDRs is difficult to measure due to significant technical challenges to survey in vivo protein conformations on a proteome-wide scale. We reasoned that IDRs should be more accessible to targeted in vivo biotinylation than more ordered protein regions, if they retain their flexibility in vivo. Indeed, we observed a positive correlation of predicted IDRs and biotinylation density across four independent large-scale proximity proteomics studies that together report >20 000 biotinylation sites. We show that biotin painting is a promising approach to fill gaps in knowledge between static in vitro protein structures, in silico disorder predictions and in vivo condition-dependent subcellular plasticity using the 80S ribosome as an example.

cell biology

DIA-NN: Deep neural networks substantially improve the identification performance of Data-independent acquisition (DIA) in proteomics

Data-independent acquisition (DIA-MS) strategies, like SWATH-MS, have been developed to increase consistency, quantification precision and proteomic depth in label-free proteomic experiments. They aim to overcome stochasticity in the selection of precursor ions by utilising (mass-) windowed acquisition that is followed by computational reconstruction of the chromatograms. While DIA methods increasingly outperform typical data-dependent methods in identification consistency and precision specifically on large sample series, possibilities remain for further improvements. At present, only a fraction of the information recorded in the complex DIA spectra is extracted by the software analysis pipelines. Here we present a software tool (DIA-NN) that introduces artificial neural nets and a new quantification strategy to enhance signal processing in DIA-data. DIA-NN greatly improves identification of precursor ions and, as a consequence, protein quantification accuracy. The performance of DIA-NN demonstrates that deep learning provides opportunities to boost the analysis of data-independent acquisition workflows in proteomics.

bioinformatics

A Bayesian Mixture Modelling Approach For Spatial Proteomics

AbstractAnalysis of the spatial sub-cellular distribution of proteins is of vital importance to fully understand context specific protein function. Some proteins can be found with a single location within a cell, but up to half of proteins may reside in multiple locations, can dynamically re-localise, or reside within an unknown functional compartment. These considerations lead to uncertainty in associating a protein to a single location. Currently, mass spectrometry (MS) based spatial proteomics relies on supervised machine learning algorithms to assign proteins to sub-cellular locations based on common gradient profiles. However, such methods fail to quantify uncertainty associated with sub-cellular class assignment. Here we reformulate the framework on which we perform statistical analysis. We propose a Bayesian generative classifier based on Gaussian mixture models to assign proteins probabilistically to sub-cellular niches, thus proteins have a probability distribution over sub-cellular locations, with Bayesian computation performed using the expectation-maximisation (EM) algorithm, as well as Markov-chain Monte-Carlo (MCMC). Our methodology allows proteome-wide uncertainty quantification, thus adding a further layer to the analysis of spatial proteomics. Our framework is flexible, allowing many different systems to be analysed and reveals new modelling opportunities for spatial proteomics. We find our methods perform competitively with current state-of-the art machine learning methods, whilst simultaneously providing more information. We highlight several examples where classification based on the support vector machine is unable to make any conclusions, while uncertainty quantification using our approach provides biologically intriguing results. To our knowledge this is the first Bayesian model of MS-based spatial proteomics data.\n\nAuthor summarySub-cellular localisation of proteins provides insights into sub-cellular biological processes. For a protein to carry out its intended function it must be localised to the correct sub-cellular environment, whether that be organelles, vesicles or any sub-cellular niche. Correct sub-cellular localisation ensures the biochemical conditions for the protein to carry out its molecular function are met, as well as being near its intended interaction partners. Therefore, mis-localisation of proteins alters cell biochemistry and can disrupt, for example, signalling pathways or inhibit the trafficking of material around the cell. The sub-cellular distribution of proteins is complicated by proteins that can reside in multiple micro-environments, or those that move dynamically within the cell. Methods that predict protein sub-cellular localisation often fail to quantify the uncertainty that arises from the complex and dynamic nature of the sub-cellular environment. Here we present a Bayesian methodology to analyse protein sub-cellular localisation. We explicitly model our data and use Bayesian inference to quantify uncertainty in our predictions. We find our method is competitive with state-of-the-art machine learning methods and additionally provides uncertainty quantification. We show that, with this additional information, we can make deeper insights into the fundamental biochemistry of the cell.

systems biology

Organisation of the transcriptional regulation of genes involved in protein transactions in yeast

The topological analyses of many large-scale molecular interaction networks often provide only limited insights into network function or evolution. In this paper, we argue that the functional heterogeneity of network components, rather than network size, is the main factor limiting the utility of topological analysis of large cellular networks. We have analysed large epistatic, functional, and transcriptional regulatory networks of genes that were attributed to the following biological process groupings: protein transactions, gene expression, cell cycle, and small molecule metabolism. Control analyses were performed on networks of randomly selected genes. We identified novel biological features emerging from the analysis of functionally homogenous biological networks irrespective of their size. In particular, direct regulation by transcription as an underrepresented feature of protein transactions. The analysis also demonstrated that the regulation of the genes involved in protein transactions at the transcriptional level was orchestrated by only a small number of regulators. Quantitative proteomic analysis of nuclear- and chromatin-enriched sub-cellular fractions of yeast provided supportive evidence for the conclusions generated by network analyses.

systems biology