Scene Dynamics: Counterfactual Critic Multi-Agent Training for Scene Graph Generation

12/06/2018
by   Long Chen, et al.
6

Scene graphs -- objects as nodes and visual relationships as edges -- describe the whereabouts and interactions of the things and stuff in an image for comprehensive scene understanding. To generate coherent scene graphs, almost all existing methods exploit the fruitful visual context by modeling message passing among objects, fitting the dynamic nature of reasoning with visual context, eg, "person" on "bike" can help determine the relationship "ride", which in turn contributes to the category confidence of the two objects. However, we argue that the scene dynamics is not properly learned by using the prevailing cross-entropy based supervised learning paradigm, which is not sensitive to graph inconsistency: errors at the hub or non-hub nodes are unfortunately penalized equally. To this end, we propose a Counterfactual critic Multi-Agent Training (CMAT) approach to resolve the mismatch. CMAT is a multi-agent policy gradient method that frames objects as cooperative agents, and then directly maximizes a graph-level metric as the reward. In particular, to assign the reward properly to each agent, CMAT uses a counterfactual baseline that disentangles the agent-specific reward by fixing the dynamics of other agents. Extensive validations on the challenging Visual Genome benchmark show that CMAT achieves a state-of-the-art by significant performance gains under various settings and metrics.

READ FULL TEXT

page 6

page 14

page 15

research
05/24/2017

Counterfactual Multi-Agent Policy Gradients

Cooperative multi-agent systems can be naturally used to model many real...
research
04/01/2020

Counterfactual Multi-Agent Reinforcement Learning with Graph Convolution Communication

We consider a fully cooperative multi-agent system where agents cooperat...
research
12/05/2018

Learning to Compose Dynamic Tree Structures for Visual Contexts

We propose to compose dynamic tree structures that place the objects in ...
research
07/18/2019

Prioritized Guidance for Efficient Multi-Agent Reinforcement Learning Exploration

Exploration efficiency is a challenging problem in multi-agent reinforce...
research
03/23/2019

An End-to-End Network for Generating Social Relationship Graphs

Socially-intelligent agents are of growing interest in artificial intell...
research
06/22/2017

Pixels to Graphs by Associative Embedding

Graphs are a useful abstraction of image content. Not only can graphs re...

Please sign up or login with your details

Forgot password? Click here to reset