bioRxiv · 10.1101/2022.04.29.490102
Statistical correction of input gradients for black box models trained with categorical input features
Abstract
Post-hoc attribution methods are widely applied to provide insights into patterns learned by deep neural networks (DNNs). Despite their success in regulatory genomics, DNNs can learn arbitrary functions outside the probabilistic simplex that defines one-hot encoded DNA. This introduces a random gradient component that manifests as noise in attribution scores. Here we demonstrate the pervasiveness of off-simplex gradient noise for genomic DNNs and introduce a statistical correction that is effective at improving the interpretability of attribution methods.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Majdandzic, A., Koo, P.. 2022-05-01. Statistical correction of input gradients for black box models trained with categorical input features. https://doi.org/10.1101/2022.04.29.490102
Cite the original work for its findings. Save a collection to share your selection of sources.