Search bioRxiv⌕ Search

Biology subjects

Malfait, M.

Publications and source records attributed to Malfait, M..

3 recordsLinked to original sources

Strategies for addressing pseudoreplication in multi-patient scRNA-seq data

The rapidly evolving field of single-cell transcriptomics has provided a powerful means for understanding cellular heterogeneity. Large-scale studies with multiple biological samples hold promise for discovering differentially expressed biomarkers with a higher level of confidence through a better characterization of the target population. However, the hierarchical nature of these experiments introduces a significant challenge for downstream statistical analysis. Indeed, despite the availability of numerous differential expression methods, only a select few can accurately address the within-patient correlation of single-cell expression profiles. Furthermore, due to the high computational costs associated with some of these methods, their practical use is limited. In this manuscript, we undertake a comprehensive assessment of different strategies to address the hierarchical correlation structure in multi-sample scRNA-seq data. We employ synthetic data generated from a simulator that retains the original correlation structure of multi-patient data while making minimal assumptions, providing a robust platform for benchmarking method performance. Our analyses indicate that neglecting within-patient correlation jeopardizes type I error control. We show that, in line with some previous reports and in contrast with others, Poisson Generalized Estimation Equations provide a useful and flexible framework for addressing these issues. We also show that pseudobulk approaches outperform single-cell level methods across the board. In this work, we resolve the conflicting results regarding the utility of GEEs and their performance relative to pseudobulk approaches. As such, we provide valuable guidelines for researchers navigating the complex landscape of gene expression modeling, and offer insights on choosing the most appropriate methods based on the specific structure and design of their datasets.

bioinformatics↗

Differential detection workflows for multi-sample single-cell RNA-seq data

In single-cell transcriptomics, differential gene expression (DE) analyses typically focus on testing differences in the average expression of genes between cell types or conditions of interest. Single-cell transcriptomics, however, also has the promise to prioritise genes for which the expression differ in other aspects of the distribution. Here we develop a workflow for assessing differential detection (DD), which tests for differences in the average fraction of samples or cells in which a gene is detected. After benchmarking eight different DD data analysis strategies, we provide a unified workflow for jointly assessing DE and DD. Using simulations and two case studies, we show that DE and DD analysis provide complementary information, both in terms of the individual genes they report and in the functional interpretation of those genes.

bioinformatics↗

Quality control for the target decoy approach for peptide identification

Reliable peptide identification is key in mass spectrometry (MS) based proteomics. To this end, the target-decoy approach (TDA) has become the cornerstone for extracting a set of reliable peptide-to-spectrum matches (PSMs) that will be used in downstream analysis. Indeed, TDA is now the default method to estimate the false discovery rate (FDR) for a given set of PSMs, and users typically view it as a universal solution for assessing the FDR in the peptide identification step. However, the TDA also relies on a minimal set of assumptions, which are typically never verified in practice. We argue that a violation of these assumptions can lead to poor FDR control, which can be detrimental to any downstream data analysis. We here therefore first clearly spell out these TDA assumptions, and introduce TargetDecoy, a Bioconductor package with all the necessary functionality to control the TDA quality and its underlying assumptions for a given set of PSMs. TOC Graphic O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=111 SRC="FIGDIR/small/516857v1_ufig1.gif" ALT="Figure 1"> View larger version (20K): org.highwire.dtl.DTLVardef@fed04borg.highwire.dtl.DTLVardef@11d115dorg.highwire.dtl.DTLVardef@15f1a45org.highwire.dtl.DTLVardef@b5ceab_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗