Kernel Single-Index Bandits: Estimation, Inference, and Learning
Sakshi Arya, Satarupa Bhattacharjee, Bharath K. Sriperumbudur
This paper studies a machine learning problem where an algorithm must repeatedly choose between different actions (arms) to maximize rewards, where each action's reward depends on input features through a single hidden direction combined with an unknown nonlinear function. The authors propose a new algorithm that can simultaneously learn which direction matters for each action, estimate the nonlinear reward function, and make statistically valid confidence intervals—all while adapting to feedback in real-time. They prove their method is both statistically efficient and achieves good cumulative performance over time.
contextual banditssingle-index modelssemiparametric inferenceonline learning