Back to papers
March 19, 2026stat.MLcs.LGmath.STAdvanced
Kernel Single-Index Bandits: Estimation, Inference, and Learning
AI-Generated Summary
This paper studies a machine learning problem where an algorithm must repeatedly choose between different actions (arms) to maximize rewards, where each action's reward depends on input features through a single hidden direction combined with an unknown nonlinear function. The authors propose a new algorithm that can simultaneously learn which direction matters for each action, estimate the nonlinear reward function, and make statistically valid confidence intervals—all while adapting to feedback in real-time. They prove their method is both statistically efficient and achieves good cumulative performance over time.
Difficulty
Advanced
Categories
stat.ML, cs.LG, math.ST
AI Tags
contextual banditssingle-index modelssemiparametric inferenceonline learningkernel methodsconfidence intervalsregret bounds