Pareto Deterministic Policy Gradients and Its Application in 5G Massive MIMO Networks

12/02/2020
by   Zhou Zhou, et al.
0

In this paper, we consider jointly optimizing cell load balance and network throughput via a reinforcement learning (RL) approach, where inter-cell handover (i.e., user association assignment) and massive MIMO antenna tilting are configured as the RL policy to learn. Our rationale behind using RL is to circumvent the challenges of analytically modeling user mobility and network dynamics. To accomplish this joint optimization, we integrate vector rewards into the RL value network and conduct RL action via a separate policy network. We name this method as Pareto deterministic policy gradients (PDPG). It is an actor-critic, model-free and deterministic policy algorithm which can handle the coupling objectives with the following two merits: 1) It solves the optimization via leveraging the degree of freedom of vector reward as opposed to choosing handcrafted scalar-reward; 2) Cross-validation over multiple policies can be significantly reduced. Accordingly, the RL enabled network behaves in a self-organized way: It learns out the underlying user mobility through measurement history to proactively operate handover and antenna tilt without environment assumptions. Our numerical evaluation demonstrates that the introduced RL method outperforms scalar-reward based approaches. Meanwhile, to be self-contained, an ideal static optimization based brute-force search solver is included as a benchmark. The comparison shows that the RL approach performs as well as this ideal strategy, though the former one is constrained with limited environment observations and lower action frequency, whereas the latter ones have full access to the user mobility. The convergence of our introduced approach is also tested under different user mobility environment based on our measurement data from a real scenario.

READ FULL TEXT

page 7

page 23

research
11/03/2021

A Self-adaptive LSAC-PID Approach based on Lyapunov Reward Shaping for Mobile Robots

To solve the coupling problem of control loops and the adaptive paramete...
research
09/28/2022

Reinforcement Learning with Tensor Networks: Application to Dynamical Large Deviations

We present a framework to integrate tensor network (TN) methods with rei...
research
05/29/2020

Reinforcement Learning

Reinforcement learning (RL) is a general framework for adaptive control,...
research
12/17/2018

User Association and Load Balancing for Massive MIMO through Deep Learning

This work investigates the use of deep learning to perform user cell ass...
research
12/27/2018

Generative Adversarial User Model for Reinforcement Learning Based Recommendation System

There are great interests as well as many challenges in applying reinfor...
research
03/13/2023

Reinforcement Learning-based Wavefront Sensorless Adaptive Optics Approaches for Satellite-to-Ground Laser Communication

Optical satellite-to-ground communication (OSGC) has the potential to im...

Please sign up or login with your details

Forgot password? Click here to reset