Signals 4 · free daily AI digest

UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training

Paper recorded by Signals 4 on 2026-09-29 in cs.AI. Abstract reproduced from arXiv; link to the original below.

Published 2026-09-29 on arXiv · recorded by Signals 4 on 2026-09-30

Category: cs.AI · 人工智能 · first seen 2026-09-30

Abstract

Interactive agent benchmarks and multi-turn reinforcement learning increasingly place a second language model in the role of the user. This simulated user controls what information the agent receives and when, yet current benchmarks score only the agent and do not directly measure whether the user correctly executed its assigned role. We introduce UserProxyBench, an evaluation layer over the tau-b

Read on arXiv →

#18 most recent of 460 cs.AI papers we have recorded · ↑ newer: Jaxolotl: A Unified High-Performance Benchmark Suite for LTL-Based Mul · ↓ older: Gender bias across LLMs is common and highly heterogenous
Cite this page: UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training: the #18 most recent of 460 cs.AI papers we have recorded (as of 2026-09-29). Source: Signals 4 (Signals API) — https://data.jiangzhang.ca/signals4/t/papers/userproxybench-evaluating-llm-user-simulators-for-agent-benchmarks-and-training.html
Free to quote with attribution to “Signals 4 (Signals API)”. Machine-readable: papers.json
Related: More cs.AI papers · arXiv signals · All papers · Today in AI
Get 4 AI signals a day by email — free.
Subscribe free → See all plans →
Get 4 AI signals a day by email — free
All models · All repos · By company · Daily editions