Signals 4 · free daily AI digest

Measuring the Serving Stack Instead of the Model: Hidden Confounds in Local Tool-Use Evaluation

Paper recorded by Signals 4 on 2026-09-22 in cs.AI. Abstract reproduced from arXiv; link to the original below.

Published 2026-09-22 on arXiv · recorded by Signals 4 on 2026-09-23

Category: cs.AI · 人工智能 · first seen 2026-09-23

Abstract

A coding agent must emit a valid tool call--a parseable invocation of a tool in the provided schema--before the harness can execute its chosen action. We study how local serving stacks affect this protocol step and show that measured outcomes can depend on the serving layer rather than model behavior alone. In Ollama, the default tools= request is gated per model by a static template flag: some mo

Read on arXiv →

#14 most recent of 360 cs.AI papers we have recorded · ↑ newer: Beyond Repeated Sampling: Learning Search Policies for LLM Reasoning · ↓ older: From Alignment to Access Control: A Framework for GenAI Policy Enforce
Cite this page: Measuring the Serving Stack Instead of the Model: Hidden Confounds in Local Tool-Use Evaluation: the #14 most recent of 360 cs.AI papers we have recorded (as of 2026-09-22). Source: Signals 4 (Signals API) — https://data.jiangzhang.ca/signals4/t/papers/measuring-the-serving-stack-instead-of-the-model-hidden-confounds-in-local-tool-.html
Free to quote with attribution to “Signals 4 (Signals API)”. Machine-readable: papers.json
Related: More cs.AI papers · arXiv signals · All papers · Today in AI
Get 4 AI signals a day by email — free.
Subscribe free → See all plans →
Get 4 AI signals a day by email — free
All models · All repos · By company · Daily editions