Near-Optimal Learning in Parametric Bandits with Action-Dependent Coarsened Feedback

To appear in Conference on Neural Information Processing Systems (NeurIPS), Dec. 2026.

Authors: Zhuohua Li, Maoli Liu, Yuwen Huang, Cheng Wen, Jie Su, Cong Tian, Shengchao Qin and John C.S. Lui

This paper studies sequential decision-making when an action determines both its reward and how much information the learner observes. It introduces a parametric bandit framework that captures the trade-off between immediate reward and informative feedback. The proposed ACF-UCB algorithm combines feedback across options to learn shared task parameters, while SW-ACF-UCB adapts to piecewise-stationary environments. Regret guarantees and lower bounds characterize the benefits of exploiting this shared structure.