Back to papers
March 19, 2026cs.HCcs.AIcs.LGIntermediate

From Accuracy to Readiness: Metrics and Benchmarks for Human-AI Decision-Making

AI-Generated Summary

This paper argues that AI systems should be evaluated not just on accuracy, but on whether humans and AI can work together safely and effectively. The authors propose a new framework with four types of metrics to measure human-AI team readiness—focusing on actual team outcomes, how much humans rely on AI, safety signals, and learning over time—rather than just testing the AI model itself.

Difficulty
Intermediate
Categories

cs.HC, cs.AI, cs.LG

AI Tags
human-AI collaborationevaluation metricsAI safetycalibrationhuman-AI teamsdecision-makingbenchmarking