Publications
CONFERENCE (INTERNATIONAL) Sample-Efficient Hypergradient Estimation for Decentralized Bi-Level Reinforcement Learning
Mikoto Kudo (University of Tsukuba/RIKEN AIP), Takumi Tanabe, Akifumi Wachi, Youhei Akimoto (University of Tsukuba/RIKEN AIP)
International Conference on Automated Planning and Scheduling (ICAPS)
June 08, 2026
Many strategic decision-making problems, such as smart traf- fic management and cybersecurity, are naturally modeled as bi-level reinforcement learning (RL), where a leader agent optimizes its objective while a follower solves a Markov de- cision process (MDP) conditioned on the leader’s decisions. In many situations, a fundamental challenge arises when the leader cannot intervene in the follower’s optimization pro- cess; it can only observe the optimization outcome. We ad- dress this decentralized setting by deriving the hypergradient of the leader’s objective, which is the gradient for the leader’s strategy that takes into account the change in the follower’s optimal policy. Unlike prior hypergradient-based methods that require extensive data for repeated state visits or rely on gradient estimators whose complexity can increase sub- stantially with the high-dimensional leader’s decision space, we leverage the Boltzmann covariance trick to derive an al- ternative hypergradient formulation. This enables efficient hypergradient estimation solely from interaction samples, even with high-dimensional leader decisions. Additionally, to our knowledge, this constitutes the first approach enabling hypergradient-based optimization for 2-player Markov games in decentralized settings. Experiments highlight the impact of hypergradient updates and demonstrate our method’s effec- tiveness in both discrete and continuous state tasks.
Paper :
Sample-Efficient Hypergradient Estimation for Decentralized Bi-Level Reinforcement Learning
(external link)