Online Learning and Equilibrium Computation with Ranking Feedback
Mingyang Liu, Yongshan Chen, Zhiyuan Fan, et al.
This paper studies online learning when the learner only receives ranking feedback (like "action A is better than B") instead of numeric scores, which is more practical for human feedback and privacy-sensitive applications. The authors show that learning with instantaneous rankings is fundamentally impossible, but develop algorithms that achieve good performance when utilities change slowly or when using time-averaged rankings. Their approach enables multiple players in games to reach approximate equilibrium through repeated play.
online learningranking feedbackgame theoryregret minimization