Risk-Sensitive Markov Decision Processes with Long-Run CVaR Criterion

10/17/2022
by   Li Xia, et al.
0

CVaR (Conditional Value at Risk) is a risk metric widely used in finance. However, dynamically optimizing CVaR is difficult since it is not a standard Markov decision process (MDP) and the principle of dynamic programming fails. In this paper, we study the infinite-horizon discrete-time MDP with a long-run CVaR criterion, from the view of sensitivity-based optimization. By introducing a pseudo CVaR metric, we derive a CVaR difference formula which quantifies the difference of long-run CVaR under any two policies. The optimality of deterministic policies is derived. We obtain a so-called Bellman local optimality equation for CVaR, which is a necessary and sufficient condition for local optimal policies and only necessary for global optimal policies. A CVaR derivative formula is also derived for providing more sensitivity information. Then we develop a policy iteration type algorithm to efficiently optimize CVaR, which is shown to converge to local optima in the mixed policy space. We further discuss some extensions including the mean-CVaR optimization and the maximization of CVaR. Finally, we conduct numerical experiments relating to portfolio management to demonstrate the main results. Our work may shed light on dynamically optimizing CVaR from a sensitivity viewpoint.

READ FULL TEXT

page 24

page 27

research
08/09/2020

Risk-Sensitive Markov Decision Processes with Combined Metrics of Mean and Variance

This paper investigates the optimization problem of an infinite stage di...
research
01/15/2022

A unified algorithm framework for mean-variance optimization in discounted Markov decision processes

This paper studies the risk-averse mean-variance optimization in infinit...
research
02/27/2023

Global Algorithms for Mean-Variance Optimization in Markov Decision Processes

Dynamic optimization of mean and variance in Markov decision processes (...
research
08/25/2019

A Complete Algebraic Transformational Solution for the Optimal Dynamic Policy in Inventory Rationing across Two Demand Classes

In this paper, we apply the sensitivity-based optimization to propose an...
research
02/17/2021

Self-Triggered Markov Decision Processes

In this paper, we study Markov Decision Processes (MDPs) with self-trigg...
research
09/01/2023

Learning Risk Preferences in Markov Decision Processes: an Application to the Fourth Down Decision in Football

For decades, National Football League (NFL) coaches' observed fourth dow...
research
04/03/2023

Investigation of risk-aware MDP and POMDP contingency management autonomy for UAS

Unmanned aircraft systems (UAS) are being increasingly adopted for vario...

Please sign up or login with your details

Forgot password? Click here to reset