Explore
Stack
Learn
Search the catalog
⌘K
$
/
₹
Sign up free
Login
Back to papers
September 4, 2026
cs.AI
Beyond Aggregate Scores: Behavioral Correctness Assumptions for Assessing Reference-Based Automatic Evaluation Methods
Maria Mahbub
,
Ashley Rice
,
Michael R. Munroe
,
Amidu Kamara
,
Amir Sadovnik
Original Abstract
Read on arXiv
Download PDF
Categories
cs.AI
Beyond Aggregate Scores: Behavioral | One9Founders