Planning in entropy-regularized Markov decision processes and games
arXiv:2604.19695v1 Announce Type: new
Abstract: We propose SmoothCruiser, a new planning algorithm for estimating the value function in entropy-regularized Markov decision processes and two-player games, given a generative model of the environment. Sm…