Best-Arm Identification with Noisy Actuation
arXiv:2604.02255v1 Announce Type: cross
Abstract: In this paper, we consider a multi-armed bandit (MAB) instance and study how to identify the best arm when arm commands are conveyed from a central learner to a distributed agent over a discrete memory…