Rethinking interpretation: Input-agnostic saliency mapping of deep visual classifiers

03/31/2023
by   Naveed Akhtar, et al.
0

Saliency methods provide post-hoc model interpretation by attributing input features to the model outputs. Current methods mainly achieve this using a single input sample, thereby failing to answer input-independent inquiries about the model. We also show that input-specific saliency mapping is intrinsically susceptible to misleading feature attribution. Current attempts to use 'general' input features for model interpretation assume access to a dataset containing those features, which biases the interpretation. Addressing the gap, we introduce a new perspective of input-agnostic saliency mapping that computationally estimates the high-level features attributed by the model to its outputs. These features are geometrically correlated, and are computed by accumulating model's gradient information with respect to an unrestricted data distribution. To compute these features, we nudge independent data points over the model loss surface towards the local minima associated by a human-understandable concept, e.g., class label for classifiers. With a systematic projection, scaling and refinement process, this information is transformed into an interpretable visualization without compromising its model-fidelity. The visualization serves as a stand-alone qualitative interpretation. With an extensive evaluation, we not only demonstrate successful visualizations for a variety of concepts for large-scale models, but also showcase an interesting utility of this new form of saliency mapping by identifying backdoor signatures in compromised classifiers.

READ FULL TEXT

page 6

page 7

research
09/08/2017

DeepFeat: A Bottom Up and Top Down Saliency Model Based on Deep Features of Convolutional Neural Nets

A deep feature based saliency model (DeepFeat) is developed to leverage ...
research
01/27/2022

Human Interpretation of Saliency-based Explanation Over Text

While a lot of research in explainable AI focuses on producing effective...
research
11/10/2020

Removing Brightness Bias in Rectified Gradients

Interpretation and improvement of deep neural networks relies on better ...
research
05/21/2018

Classifier-agnostic saliency map extraction

We argue for the importance of decoupling saliency map extraction from a...
research
05/13/2021

Sanity Simulations for Saliency Methods

Saliency methods are a popular class of feature attribution tools that a...
research
04/05/2018

End-to-End Saliency Mapping via Probability Distribution Prediction

Most saliency estimation methods aim to explicitly model low-level consp...
research
10/20/2022

Finding Dataset Shortcuts with Grammar Induction

Many NLP datasets have been found to contain shortcuts: simple decision ...

Please sign up or login with your details

Forgot password? Click here to reset