Search bioRxiv⌕ Search

Biology subjects

Khang, T. F.

Publications and source records attributed to Khang, T. F..

2 recordsLinked to original sources

SIEVE: One-stop differential expression, variability, and skewness analyses using RNA-Seq data

RNA-Seq data analysis is commonly biased towards detecting differentially expressed genes and insufficiently conveys the complexity of gene expression changes between biological conditions. This bias arises because discrete count models cannot fully and independently parameterize the mean, variance, and skewness of gene expression distributions. Therefore, a unified statistical framework that simultaneously tests differential expression, variability, and skewness is needed. We present SIEVEseq, a statistical methodology that provides such a framework. SIEVEseq embraces a compositional data analysis strategy to transform discrete RNA-Seq counts into continuous form with a distribution well-fitted by the skew-normal distribution. Both parametric and nonparametric simulations show that SIEVEseq better controls the false discovery rate and Type II error than existing differential expression methods. Analysis of the Mayo RNA-Seq dataset for Alzheimers disease demonstrates that gene sets with significant differences in mean, variance, and skewness between control and disease groups strongly predict disease state. Furthermore, functional enrichment analysis indicates that relying solely on differentially expressed genes identifies only part of the biological spectrum, whereas incorporating genes with differential variability and skewness reveals additional disease-related aspects. Cross-data and cross-methodology validation suggest the detected biological signals are genuine. The SIEVEseq R package and source codes are available at: https://github.com/Divo-Lee/SIEVEseq.

bioinformatics↗

clrDV: A differential variability test for RNA-Seqdata based on the skew-normal distribution

Genes that show differential variability between conditions are important for complementing a systems biology understanding of the molecular players involved in a biological process. Under the dominant paradigm for modeling RNA-Seq gene counts using the negative binomial model, tests of differential variability are challenging to develop, owing to dependence of the variance on the mean. The limited availability of methods for detecting genes with differential variability means that researchers often omit differential variability as an analytical step in RNA-Seq data analysis. Here, we describe clrDV, a statistical method for detecting genes that show differential variability between two populations. clrDV is based on a compositional data analysis framework. We present the skew-normal distribution for modeling gene-wise null distribution of centered log-ratio transformation of compositional RNA-seq data. Simulation results show that clrDV has false discovery rate and Type II error that are on par with or superior to existing methodologies. In addition, its run time is faster than the closest competitors, and remains relatively constant for increasing sample size per group. Analysis of a large neurodegenerative disease RNA-Seq dataset using clrDV recovers multiple gene candidates that have been reported to be associated with Alzheimers disease. Additionally, we find that the majority of genes with differential variability have smaller relative gene expression variance in the Alzheimers disease population compared to the control population.

bioinformatics↗