Search bioRxiv⌕ Search

Biology subjects

Skowronek, P.

Publications and source records attributed to Skowronek, P..

4 recordsLinked to original sources

AlphaViz: Visualization and validation of critical proteomics data directly at the raw data level

Although current mass spectrometry (MS)-based proteomics identifies and quantifies thousands of proteins and (modified) peptides, only a minority of them are subjected to in-depth downstream analysis. With the advent of automated processing workflows, biologically or clinically important results within a study are rarely validated by visualization of the underlying raw information. Current tools are often not integrated into the overall analysis nor readily extendable with new approaches. To remedy this, we developed AlphaViz, an open-source Python package to superimpose output from common analysis workflows on the raw data for easy visualization and validation of protein and peptide identifications. AlphaViz takes advantage of recent breakthroughs in the deep learning-assisted prediction of experimental peptide properties to allow manual assessment of the expected versus measured peptide result. We focused on the visualization of the 4-dimensional data cuboid provided by Bruker TimsTOF instruments, where the ion mobility dimension, besides intensity and retention time, can be predicted and used for verification. We illustrate how AlphaViz can quickly validate or invalidate peptide identifications regardless of the score given to them by automated workflows. Furthermore, we provide a predict mode that can locate peptides present in the raw data but not reported by the search engine. This is illustrated the recovery of missing values from experimental replicates. Applied to phosphoproteomics, we show how key signaling nodes can be validated to enhance confidence for downstream interpretation or follow-up experiments. AlphaViz follows standards for open-source software development and features an easy-to-install graphical user interface for end-users and a modular Python package for bioinformaticians. Validation of critical proteomics results should now become a standard feature in MS-based proteomics.

bioinformatics↗

Rapid and in-depth coverage of the (phospho-)proteome with deep libraries and optimal window design for dia-PASEF

Data-independent acquisition (DIA) methods have become increasingly attractive in mass spectrometry (MS)-based proteomics, because they enable high data completeness and a wide dynamic range. Recently, we combined DIA with parallel accumulation - serial fragmentation (dia-PASEF) on a Bruker trapped ion mobility separated (TIMS) quadrupole time-of-flight (TOF) mass spectrometer. This requires alignment of the ion mobility separation with the downstream mass selective quadrupole, leading to a more complex scheme for dia-PASEF window placement compared to DIA. To achieve high data completeness and deep proteome coverage, here we employ variable isolation windows that are placed optimally depending on precursor density in the m/z and ion mobility plane. This Automatic Isolation Design procedure is implemented in the freely available py_diAID package. In combination with in-depth project-specific proteomics libraries and the Evosep LC system, we reproducibly identified over 7,700 proteins in a human cancer cell line in 44 minutes with quadruplicate single-shot injections at high sensitivity. Even at a throughput of 100 samples per day (11 minutes LC gradients), we consistently quantified more than 6,000 proteins in mammalian cell lysates by injecting four replicates. We found that optimal dia-PASEF window placement facilitates in-depth phosphoproteomics with very high sensitivity, quantifying more than 35,000 phosphosites in a human cancer cell line stimulated with an epidermal growth factor (EGF) in triplicate 21 minutes runs. This covers a substantial part of the regulated phosphoproteome with high sensitivity, opening up for extensive systems-biological studies.

biochemistry↗

AlphaTims: Indexing trapped ion mobility spectrometry - time of flight data for fast and easy accession and visualization

High resolution mass spectrometry-based proteomics generates large amounts of data, even in the standard liquid chromatography (LC) - tandem mass spectrometry configuration. Adding an ion mobility dimension vastly increases the acquired data volume, challenging both analytical processing pipelines and especially data exploration by scientists. This has necessitated data aggregation, effectively discarding much of the information present in these rich data sets. Taking trapped ion mobility spectrometry (TIMS) on a quadrupole time-of-flight platform (Q-TOF) as an example, we developed an efficient indexing scheme that represents all data points as detector arrival times on scales of minutes (LC), milliseconds (TIMS) and microseconds (TOF). In our open source AlphaTims package, data are indexed, accessed and visualized by a combination of tools of the scientific Python ecosystem. We interpret unprocessed data as a sparse 4D matrix and use just-in-time compilation to machine code with Numba, accelerating our computational procedures by several orders of magnitude while keeping to familiar indexing and slicing notations. For samples with more than six billion detector events, a modern laptop can load and index raw data in about a minute. Loading is even faster when AlphaTims has already saved indexed data in a HDF5 file, a portable scientific standard used in extremely large-scale data acquisition. Subsequently, data accession along any dimension and interactive visualization happen in milliseconds. We have found AlphaTims to be a key enabling tool to explore high dimensional LC-TIMS-QTOF data and have made it freely available as an open-source Python package with a stand-alone graphical user interface at https://github.com/MannLabs/alphatims or as part of the AlphaPept ecosystem. HighlightsO_LIEasy visualization and fast accession of LC-TIMS-QTOF data C_LIO_LIFreely available graphical user interface, command-line interface and Python module on Windows, Linux and macOS. C_LI

bioinformatics↗

Residue-specific insights into (2x)72 kDa tryptophan synthase obtained from fast-MAS 1H-detected solid-state NMR

Solid-state NMR has emerged as a potent technique in structural biology, suitable for the study of fibrillar, micro-crystalline, and membrane proteins. Recent developments in fast-magic-angle-spinning and proton-detected methods have enabled detailed insights into structure and dynamics, but molecular-weight limitations for the asymmetric part of target proteins have remained at ~30-40 kDa. Here we employ solid-state NMR for atom-specific characterization of the 72 kDa (asymmetric unit) microcrystalline protein tryptophan synthase, an important target in pharmacology and biotechnology, chemical-shift assignments of which we obtain via higher-dimensionality, 4D and 5D solid-state NMR experiments. The assignments for the first time provide comprehensive data for assessment of side chain chemical properties involved in the catalytic turnover, and, in conjunction with first-principles calculations, precise determination of thermodynamic and kinetic parameters is demonstrated for the essential acid-base catalytic residue {beta}K87. The insights provided by this study expand by nearly a factor of two the size limitations widely accepted for NMR today, demonstrating the applicability of solid-state NMR to systems that have been thought to be out of reach due to their complexity.

biochemistry↗