Evaluating Counterfactual Strategic Reasoning in Large Language Models
Dimitrios Georgousis, Maria Lymperaiou, Angeliki Dimitriou, et al.
Researchers tested whether large language models can genuinely reason about strategy in games or if they're just recalling familiar patterns. They evaluated LLMs in modified versions of classic games like Prisoner's Dilemma and Rock-Paper-Scissors with altered rules and rewards, finding that the models struggle to adapt their strategies to these new scenarios and don't truly understand the underlying incentives.
game theorystrategic reasoningLLM evaluationgeneralization