Decentralized partially observable Markov decision processes (Dec-POMDPs) are becoming increasingly popular in various applications ranging from decentralized control of fleet of autonomous vehicles to that of smart grids. Optimally solving Dec-POMDPs is notoriously hard as is illustrated by the non-stationary problem and the search complexity of finding best history based policies (which is NEXP complete). Agent-state based policies have emerged as a popular paradigm to address some of these challenges. In this talk, we review the existing solution approaches to find optimal agent state base policies and present a novel policy search algorithm which has monotonic improvement guarantee and converges to a locally optimal solution. We conclude by presenting experimental results that show that that the proposed algorithm identifies close to optimal policies in various POMDP and Dec-POMDP benchmarks. Joint work with Amit Sinha and Matthieu Geist.
Aditya Mahajan is Professor of Electrical and Computer Engineering at McGill University, Montreal, Canada. He is a member of the McGill Center of Intelligent Machines (CIM), Mila - Québec AI Institute, International Laboratory for Learning Systems (ILLS), and Groupe d’études et de recherche en analyse des décisions (GERAD). He received the B.Tech degree in Electrical Engineering from the Indian Institute of Technology, Kanpur, India and the MS and PhD degrees in Electrical Engineering and Computer Science from the University of Michigan, Ann Arbor, USA. He has held visiting appointments at the University of California, Berkeley and the University of Paris-Saclay. He is a senior member of the IEEE and member of Professional Engineers Ontario. He currently serves as Associate Editor of Springer Mathematics of Control, Signal, and Systems. In the past, he has served as an Associate Editor of IEEE Transactions on Automatic Control IEEE Control Systems Letters, and IEEE Control Systems Society Conference Editorial Board. He is the recipient of the 2015 George Axelby Outstanding Paper Award, the 2016 NSERC Discovery Accelerator Award, the 2014 CDC Best Student Paper Award (as supervisor), and the 2016 NecSys Best Student Paper Award (as supervisor). His principal research interests include decentralized stochastic control, team theory, reinforcement learning, multi-armed bandits and information theory.