DeepAI
Log In Sign Up

Learning Explainable Models Using Attribution Priors

06/25/2019
by   Gabriel Erion, et al.
10

Two important topics in deep learning both involve incorporating humans into the modeling process: Model priors transfer information from humans to a model by constraining the model's parameters; Model attributions transfer information from a model to humans by explaining the model's behavior. We propose connecting these topics with attribution priors (https://github.com/suinleelab/attributionpriors), which allow humans to use the common language of attributions to enforce prior expectations about a model's behavior during training. We develop a differentiable axiomatic feature attribution method called expected gradients and show how to directly regularize these attributions during training. We demonstrate the broad applicability of attribution priors (Ω) by presenting three distinct examples that regularize models to behave more intuitively in three different domains: 1) on image data, Ω_pixel encourages models to have piecewise smooth attribution maps; 2) on gene expression data, Ω_graph encourages models to treat functionally related genes similarly; 3) on a health care dataset, Ω_sparse encourages models to rely on fewer features. In all three domains, attribution priors produce models with more intuitive behavior and better generalization performance by encoding constraints that would otherwise be very difficult to encode using standard model priors.

READ FULL TEXT

page 7

page 16

page 20

page 21

page 22

page 23

12/20/2019

Learned Feature Attribution Priors

Deep learning models have achieved breakthrough successes in domains whe...
06/19/2019

Incorporating Priors with Feature Attribution on Text Classification

Feature attribution methods, proposed recently, help users interpret the...
10/15/2021

Combining Diverse Feature Priors

To improve model generalization, model designers often restrict the feat...
11/15/2021

Fast Axiomatic Attribution for Neural Networks

Mitigating the dependence on spurious correlations present in the traini...
06/15/2021

Keep CALM and Improve Visual Feature Attribution

The class activation mapping, or CAM, has been the cornerstone of featur...
05/31/2021

The effectiveness of feature attribution methods and its correlation with automatic evaluation scores

Explaining the decisions of an Artificial Intelligence (AI) model is inc...
03/18/2022

Transferable Class-Modelling for Decentralized Source Attribution of GAN-Generated Images

GAN-generated deepfakes as a genre of digital images are gaining ground ...

Code Repositories

path_explain

A repository for explaining feature attributions and feature interactions in deep neural networks.


view repo

attributionpriors

Tools for training explainable models using attribution priors.


view repo