Paper recorded by Signals 4 on 2026-09-17 in cs.AI. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-17 on arXiv · recorded by Signals 4 on 2026-09-18
Category: cs.AI · 人工智能 · first seen 2026-09-18
Evaluating an AI system requires disaggregated assessment, as performance varies across domains such as benchmark task types or conversation types in deployed agents. Exhaustive testing is expensive, so evaluation rests on a sample of labeled units. We treat the evaluation set as a finite population and seek accurate point and interval estimates of each domain mean. Direct estimators, including pr