Near-Optimal Learning in Parametric Bandits with Action-Dependent Coarsened Feedback
To appear in Conference on Neural Information Processing Systems (NeurIPS), Dec. 2026.
Authors: Zhuohua Li, Maoli Liu, Yuwen Huang, Cheng Wen, Jie Su, Cong Tian, Shengchao Qin and John C.S. Lui
This paper studies sequential decision-making when an action determines both its reward and how much information the learner observes. It introduces a parametric bandit framework that captures the trade-off between immediate reward and informative feedback. The proposed ACF-UCB algorithm combines feedback across options to learn shared task parameters, while SW-ACF-UCB adapts to piecewise-stationary environments. Regret guarantees and lower bounds characterize the benefits of exploiting this shared structure.
