Novel Data Transformations for RNA-seq Data Analysis
We propose eight data transformations for RNA-seq data analysis aiming to make the transformed sample mean to be representative of the distribution center since it is not always possible to transform count data to satisfy the normality assumption. Simulation studies showed that limma based on transformed data by using the rv transformation (denoted as limma+rv) performed best compared with limma based on transformed data by using other transformation methods in term of high accuracy and low FNR, while keeping FDR at the nominal level. For large sample size, limma based on transformed data by using the 8 proposed transformation methods had similar performance to limma based on transformed data by using existing transformation methods for equal library size scenarios. Otherwise, limma based on transformed data by using the rv, lv, rv2, or lv2 transformation, or by using the existing voom transformation performed better than limma based on data from other transformation methods. Real data analysis results showed that limma+ l2 performed best, while limma+ rv also had good performance.