Logo image
On the Convergence of Modified Policy Iteration in Risk-Sensitive Exponential Cost Markov Decision Processes
Journal article   Peer reviewed

On the Convergence of Modified Policy Iteration in Risk-Sensitive Exponential Cost Markov Decision Processes

Yashaswini Murthy, Mehrdad Moharrami and Rayadurgam Srikant
Operations research, Vol.74(3), pp.1425-1436
05/2026
DOI: 10.1287/opre.2024.0818

View Online

Abstract

Modified policy iteration (MPI) is a dynamic programming algorithm that combines elements of policy iteration and value iteration. The convergence of MPI is wellstudied in the context of discounted and average-cost Markov decision processes (MDPs). In this work, we consider the exponential cost risk-sensitive MDP formulation, which is known to provide some robustness to model parameters. Although policy iteration and value iteration are well-studied in the context of risk-sensitive MDPs, MPI is unexplored. To the best of our knowledge, we provide the first proof that MPI also converges for the risk-sensitive problem in the case of finite state and action spaces. Because the exponential cost formulation deals with the multiplicative Bellman equation, our main contribution is a convergence proof that is quite different than existing results for discounted and risk-neutral average-cost as well as risk-sensitive value and iteration
Social Sciences Technology Business & Economics Management Operations Research & Management Science Science & Technology

Details

Metrics

26 Record Views
Logo image