Abdelghani Ghanem, Mounir Ghogho

Entropy-Regularized Adjoint Matching for Offline RL

Abdelghani Ghanem, Mounir Ghogho / May 8, 2026

arXiv:2605.06156v1 Announce Type: cross
Abstract: Integrating expressive generative policies, such as flow-matching models, into offline reinforcement learning (RL) allows agents to capture complex, multi-modal behaviors. While Q-learning with Adjoint…

Author name: Abdelghani Ghanem, Mounir Ghogho

Entropy-Regularized Adjoint Matching for Offline RL