Search bioRxiv⌕ Search

Biology subjects

Koh, H.

Publications and source records attributed to Koh, H..

2 recordsLinked to original sources

A general kernel machine comparative analysis framework for randomized block designs

MotivationThere are numerous potential confounders, including genetic, environmental, technical, and demographic factors. These factors may be known or unknown, measured or unmeasured; hence, it is extremely challenging to capture them in downstream data analysis. However, randomized block design is an efficient design technique to control confounding factors and to reduce variability within subjects. This helps prevent spurious discoveries and boost test power. I also note that kernel machine comparative analysis is widely employed in high-dimensional omics studies to boost test power by combining possibly weak effects from multiple underlying variants, and also to explore various linear or nonlinear patterns of disparity. ResultsIn this paper, I introduce a general kernel machine comparative analysis framework for randomized block designs, named as KernRBD, to investigate the effects of treatments (e.g., medical treatment, environmental exposure) on the underlying variants. KernRBD is unique in its range of functionalities, including the computation of P-value for global testing and adjusted P-values for pairwise comparisons, as well as visual representation through ordination plotting. KernRBD is practical, requiring only a kernel as input, and also robustly valid based on a resampling scheme not requiring the assumption of normality to be satisfied. I also introduce its omnibus test for a unified and powerful significance testing across multiple input kernels. While its applications should be much broader, I illustrate its use through human microbiome {beta}-diversity analysis in praxis, and its outperformance in significance testing through simulation experiments in silico. Availability and ImplementationKernRBD is available at https://github.com/hk1785/kernrbd.

bioinformatics↗

An ensemble learning method for joint kernel association testing and principal component analysis on multiple kernels

In high-dimensional omics studies, researchers often conduct kernel association testing to power-fully detect the relationship of the genetic or microbial composition with human health or disease. Especially, in human microbiome studies, its dimension reduction analysis follows to visually represent complex microbiome data in a simple two- or three-dimensional coordinate space. However, various kernels exist, and they produce all different outcomes; hence, it is hard to interpret them all consistently. Then, omnibus testing has recently been a subject of intense investigation for a unified and powerful statistical inference. However, current omnibus tests are purely a test for significance producing only a P-value as their outcome with no related dimension reduction and visualization approach; hence, their utility is still limited. In this paper, I introduce an ensemble learning method, named as enKern, for joint kernel associating testing and principal component analysis on multiple kernels. enKern is based on a weight learning scheme that leverages complementary contributions from multiple kernels for powerful performance for various association patterns. I show that applying the weights to individual test statistics or individual kernels is equivalent, which in turn enables a visualization in a reduced dimensional coordinate space based on the weighted kernel to be matched with its original significance testing scheme. I demonstrate its use for human microbiome {beta}-diversity analysis. I also demonstrate its outperformance in validity and power through simulation experiments. enKern is freely available in R computing environment at https://github.com/hk1785/enkern.

bioinformatics↗