bioRxiv · 10.1101/2023.12.04.570016
Normalization of RNA-Seq Data using Adaptive Trimmed Mean with Multi-reference
Abstract
The normalization of RNA sequencing data is a primary step for downstream analysis. The most popular method used for the normalization is the trimmed mean of M values (TMM) and DESeq. The TMM tries to trim away extreme log fold changes of the data to normalize the raw read counts based on the remaining non-deferentially expressed genes. However, the major problem with the TMM is that the values of trimming factor M are heuristic. This paper tries to estimate the adaptive value of M in TMM based on Jaeckels Estimator, and each sample acts as a reference to find the scale factor of each sample. The presented approach is validated on SEQC, MAQC2, MAQC3, PICKRELL, and two simulated datasets with two groups and three groups conditions by varying the percentage of differential expression and the number of replicates. The performance of the present approach is compared, and it shows better in terms of area under the receiver operating characteristic curve (AUC) and differential expression. The implementation of the present approach is available on the GitHub platform: https://github.com/vikkyak/Normalization-of-Bulk-RNA-seq.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Singh, V., Lee, S., Kirtipal, N., Song, B.-S.. 2023-12-07. Normalization of RNA-Seq Data using Adaptive Trimmed Mean with Multi-reference. https://doi.org/10.1101/2023.12.04.570016
Cite the original work for its findings. Save a collection to share your selection of sources.