PTDE: Personalized Training with Distillated Execution for Multi-Agent Reinforcement Learning

10/17/2022
by   Yiqun Chen, et al.
0

Centralized Training with Decentralized Execution (CTDE) has been a very popular paradigm for multi-agent reinforcement learning. One of its main features is making full use of the global information to learn a better joint Q-function or centralized critic. In this paper, we in turn explore how to leverage the global information to directly learn a better individual Q-function or individual actor. We find that applying the same global information to all agents indiscriminately is not enough for good performance, and thus propose to specify the global information for each agent to obtain agent-specific global information for better performance. Furthermore, we distill such agent-specific global information into the agent's local information, which is used during decentralized execution without too much performance degradation. We call this new paradigm Personalized Training with Distillated Execution (PTDE). PTDE can be easily combined with many state-of-the-art algorithms to further improve their performance, which is verified in both SMAC and Google Research Football scenarios.

READ FULL TEXT
research
05/27/2023

Is Centralized Training with Decentralized Execution Framework Centralized Enough for MARL?

Centralized Training with Decentralized Execution (CTDE) has recently em...
research
05/05/2022

General sum stochastic games with networked information flows

Inspired by applications such as supply chain management, epidemics, and...
research
05/25/2022

Scalable Multi-Agent Model-Based Reinforcement Learning

Recent Multi-Agent Reinforcement Learning (MARL) literature has been lar...
research
09/19/2021

Regularize! Don't Mix: Multi-Agent Reinforcement Learning without Explicit Centralized Structures

We propose using regularization for Multi-Agent Reinforcement Learning r...
research
04/25/2023

SEA: A Spatially Explicit Architecture for Multi-Agent Reinforcement Learning

Spatial information is essential in various fields. How to explicitly mo...
research
01/03/2022

A Deeper Understanding of State-Based Critics in Multi-Agent Reinforcement Learning

Centralized Training for Decentralized Execution, where training is done...
research
02/24/2021

Credit Assignment with Meta-Policy Gradient for Multi-Agent Reinforcement Learning

Reward decomposition is a critical problem in centralized training with ...

Please sign up or login with your details

Forgot password? Click here to reset