Back to papers
March 19, 2026cs.CLIntermediate
RADIUS: Ranking, Distribution, and Significance - A Comprehensive Alignment Suite for Survey Simulation
AI-Generated Summary
This paper addresses the problem of inconsistent evaluation methods for AI survey simulations by introducing RADIUS, a standardized evaluation framework that measures both whether simulated responses match human preferences (ranking alignment) and whether the overall distribution of responses is accurate (distribution alignment), along with statistical significance testing. The work shows that existing metrics can be misleading—a simulation might report high accuracy while still getting people's top choices wrong, which matters for real-world decision-making. RADIUS provides researchers with a unified, open-source tool to consistently compare different survey simulation approaches.
Difficulty
Intermediate
Categories
cs.CL
AI Tags
LLMsevaluation metricssurvey simulationalignmentbenchmarking