Investigating Effects of Saturation in Integrated Gradients

Human Interpretability Workshop at ICML


Integrated gradients is a popular method for post-hoc model interpretability. Despite its popularity, the composition and relative impact of different regions of the integral are not well understood. We explore these effects and find that gradients in saturated regions of the scaling factor, where model output changes minimally, contribute disproportionately to the computed attribution. We propose a variant of Integrated Gradients which primarily captures gradients in unsaturated regions and evaluate this method on ImageNet classification networks. We find that this attribution technique shows higher model faithfulness and lower sensitivity to noise than standard Integrated Gradients.

Related Publications

All Publications

NeurIPS - December 5, 2021

Local Differential Privacy for Regret Minimization in Reinforcement Learning

Evrard Garcelon, Vianney Perchet, Ciara Pike-Burke, Matteo Pirotta

NeurIPS - December 5, 2021

Hierarchical Skills for Efficient Exploration

Jonas Gehring, Gabriel Synnaeve, Andreas Krause, Nicolas Usunier

NeurIPS - December 5, 2021

Interpretable agent communication from scratch (with a generic visual processor emerging on the side)

Roberto Dessì, Eugene Kharitonov, Marco Baroni

Workshop on Online Abuse and Harms (WHOAH) at ACL - November 30, 2021

Findings of the WOAH 5 Shared Task on Fine Grained Hateful Memes Detection

Lambert Mathias, Shaoliang Nie, Bertie Vidgen, Aida Davani, Zeerak Waseem, Douwe Kiela, Vinodkumar Prabhakaran

