Search bioRxiv⌕ Search

Biology subjects

Allendes Osorio, R. S.

Publications and source records attributed to Allendes Osorio, R. S..

2 recordsLinked to original sources

OddSNP: a predictive framework for optimizing multiplexed single-cell RNA-seq experiments

Donor multiplexing is a powerful strategy to increase scale, lower the costs, and reduce batch effects in single-cell RNA sequencing (scRNAseq), but clear guidelines for experimental design are lacking, forcing researchers to risk costly demultiplexing failures. To address this, we introduce SNP-Information Content (SNP-IC), a quantitative metric computable from simple unpooled pilot data that accurately predicts the success of genotype-based demultiplexing. Across multiple large-scale datasets using stem cell and organoid models, we establish a robust SNP-IC threshold of approximately 50, above which cells can be reliably assigned to their donor of origin. For more challenging genotype-free approaches, we define a pairwise metric, cpSNP-IC, and demonstrate a much higher requirement of approximately 3,000. Our open-source framework, oddSNP, implements this predictive model, allowing researchers to perform in silico titrations of sequencing depth and donor complexity to optimize experimental design before committing to large-scale studies. oddSNP provides a practical framework, enabling researchers to strategically optimize sequencing depth and donor numbers to maximize experimental success while managing costs and minimizing the risk of catastrophic data loss.

bioinformatics↗

Snaq: A Dynamic Snakemake Pipeline for Microbiome data analysis with QIIME2

Optimizing a protocol for 16S microbiome data analysis with QIIME2 is a challenging task for biologists with no programming experience. It involves a multi-step process, and multiple parameters and options that need to be tested and determined. In this article, we describe Snaq, a snakemake pipeline that helps automate and optimize 16S data analysis using QIIME2. Snaq offers an informative file naming system and automatically performs the analysis of a data set by downloading and installing the required databases and classifiers, all through a single command-line instruction. It works natively on Linux and Mac and on Windows through the use of containers, and is potentially extendable by adding new rules. This pipeline will substantially reduce the efforts in sending commands and prevent the confusion caused by the accumulation of analysis results due to testing multiple parameters.

bioinformatics↗