OS-Themis: A Scalable Critic Framework for Generalist GUI Rewards
Zehao Li, Zhenyu Wu, Yibo Zhao, et al.
OS-Themis is a new framework that helps train AI agents to better interact with graphical user interfaces (like phone apps) by providing more reliable feedback on whether the agent is performing tasks correctly. Instead of using a single judge, it breaks down agent actions into verifiable milestones and cross-checks the evidence before making a decision, similar to how a court system works. When tested on smartphone tasks, this approach improved performance by about 10% during training and 7% when filtering practice data.
reinforcement learningGUI agentsreward functionsmulti-agent systems