RADIUS: Ranking, Distribution, and Significance - A Comprehensive Alignment Suite for Survey Simulation
Weronika Łajewska, Paul Missault, George Davidson, et al.
This paper addresses the problem of inconsistent evaluation methods for AI survey simulations by introducing RADIUS, a standardized evaluation framework that measures both whether simulated responses match human preferences (ranking alignment) and whether the overall distribution of responses is accurate (distribution alignment), along with statistical significance testing. The work shows that existing metrics can be misleading—a simulation might report high accuracy while still getting people's top choices wrong, which matters for real-world decision-making. RADIUS provides researchers with a unified, open-source tool to consistently compare different survey simulation approaches.
LLMsevaluation metricssurvey simulationalignment