iDEG: A single-subject method for assessing gene differential expression from two transcriptomes of an individual
BackgroundAccurate profiling of gene expression in a single subject has the potential to be a powerful precision medicine tool, useful for unveiling individual disease mechanisms and responses. However, expression analysis tools for RNA-sequencing (RNA-Seq) data require replicate samples to estimate gene-wise data variability and make inferences, which is costly and not easily obtainable in clinical practice. Strategies to implement DEGSeq, DESeq, and edgeR for comparing two conditions without replicates (TCWR) have been proposed without evaluation, while NOISeq-sim was validated in a restricted way using qPCR on 400 transcripts. These methods impose restrictive assumptions in TCWR limiting inferential opportunities.\n\nMethodsWe propose a new method that borrows information across different genes from the same individual using a partitioned window to strategically bypass the requirement of replicates per condition. We termed this method \"iDEG\", which identifies individualized Differentially Expressed Genes in a single subject sampled under two conditions without replicates, i.e., a baseline sample (unaffected tissue) vs. a case sample (tumor). iDEG transforms RNA-Seq data such that, under the null hypothesis, differences of transformed expression counts follow a distribution and variance calculated across a local partition of related transcripts at baseline expression. This transformation enables modeling genes with a two-group mixture model from which the probability of differential expression for each gene is then estimated by an empirical Bayes approach with a local false discovery rate control. To compare the performance of iDEG to other methods applied to TCWR, we conducted simulations assuming a Negative Binomial distribution with varying dispersion parameters and percentages of differentially expressed genes (DEGs).\n\nResultsOur extensive simulation studies demonstrate that iDEGs F1 accuracy scores better than the other methods at 5% 90% and recall>75% and low false positive rate (<1%) in most conditions.\n\nConclusionThe partitioned window strategy provides a novel and accurate way to borrow information across genes locally and would probably increase the accuracy of all relevant methods.