Paper recorded by Signals 4 on 2026-09-16 in cs.AI. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-16 on arXiv · recorded by Signals 4 on 2026-09-17
Category: cs.AI · 人工智能 · first seen 2026-09-17
Agent benchmarks are substantially more costly to evaluate than conventional LLM benchmarks. Benchmark compression is therefore a natural solution, yet existing methods primarily model redundancy in task--model final-score distributions, which is important in agentic evaluation. To address this limitation, we analyze large-scale trajectories and identify six complementary process signals that are