Back to papers
May 7, 2026cs.LGcs.AI

On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR

Categories

cs.LG, cs.AI