Search bioRxiv⌕ Search

Biology subjects

McGovern, K. C.

Publications and source records attributed to McGovern, K. C..

2 recordsLinked to original sources

Replacing Normalizations with Interval AssumptionsImproves the Rigor and Robustness of DifferentialExpression and Differential Abundance Analyses

Standard methods for differential expression and differential abundance analysis rely on normalization to address sample-to-sample variation in sequencing depth. However, normalizations imply strict, unrealistic assumptions about the unmeasured scale of biological systems (e.g., microbial load or total cellular transcription). This introduces bias that can lead to false positives and false negatives. To overcome these limitations, we suggest replacing normalizations with interval assumptions. This approach allows researchers to explicitly define plausible lower and upper bounds on the unmeasured biological systems scale, making these assumptions more realistic, transparent, and flexible than those imposed by traditional normalizations. Compared to recent alternatives like scale models and sensitivity analyses, interval assumptions are easier to use, resulting in potentially reduced false positives and false negatives, and have stronger guarantees of Type-I error control. We make interval assumptions accessible by introducing a modified version of ALDEx2 as a publicly available software package. Through simulations and real data studies, we show these methods can reduce false positives and false negatives compared to normalization-based tools.

bioinformatics↗

Addressing Erroneous Scale Assumptions in Microbe and Gene Set Enrichment Analysis

By applying Differential Set Analysis (DSA) to sequence count data, researchers can determine whether groups of microbes or genes are differentially enriched. Yet these data lack information about the scale (i.e., size) of the biological system under study, leading some authors to call these data compositional (i.e., proportional). In this article we show that commonly used DSA methods make strong, implicit assumptions about the unmeasured system scale. We show that even small errors in these assumptions can lead to false positive rates as high as 70%. To mitigate this problem, we introduce a sensitivity analysis framework to identify when modeling results are robust to such errors and when they are suspect. Unlike standard benchmarking studies, our methods do not require ground-truth knowledge and can therefore be applied to both simulated and real data.

bioinformatics↗