PACE: Perceived-Latency-Aware Cascading Service Routing and Filler Control for QoE-Efficient Retrieval-Augmented Dialogue Serving
Paper recorded by Signals 4 on 2026-09-09 in cs.AI. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-09 on arXiv · recorded by Signals 4 on 2026-09-10
Category: cs.AI · 人工智能 · first seen 2026-09-10
Abstract
We present the PACE, a framework for retrieval-augmented dialogue serving that formalizes Perceived Time-to-First-Response (PTFR) as a QoE objective and minimizes it under quality/cost constraints. Unlike prior work on cascaded routing, semantic caching, or adaptive retrieval, PACE jointly controls which answer source composes the response and what fills the waiting window. Deployed on a humanoid-
Read on arXiv →
Cite this page: PACE: Perceived-Latency-Aware Cascading Service Routing and Filler Control for QoE-Efficient Retrieval-Augmented Dialogue Serving: the #131 most recent of 300 cs.AI papers we have recorded (as of 2026-09-09). Source: Signals 4 (Signals API) — https://data.jiangzhang.ca/signals4/t/papers/pace-perceived-latency-aware-cascading-service-routing-and-filler-control-for-qo.html
Free to quote with attribution to “Signals 4 (Signals API)”. Machine-readable:
papers.json
Get 4 AI signals a day by email — free.
Get 4 AI signals a day by email — free