Signals 4 · free daily AI digest

LongHarness Bench: Stress-Testing Language Model Harnesses for Long-Context Reasoning

Paper recorded by Signals 4 on 2026-09-29 in cs.CL. Abstract reproduced from arXiv; link to the original below.

Published 2026-09-29 on arXiv · recorded by Signals 4 on 2026-09-30

Category: cs.CL · 自然语言处理 · first seen 2026-09-30

Abstract

Language-model (LM) harnesses enable LMs to operate effectively over long contexts using additional compute. However, existing long-context evaluations are insufficient for distinguishing modern harnesses, reflected by saturated accuracy across harnesses and largely similar evaluation costs. In this paper, we introduce a benchmark for evaluating both the effectiveness and efficiency of long-contex

Read on arXiv →

#4 most recent of 292 cs.CL papers we have recorded · ↑ newer: Pretraining Latent Information Feedback Transformers with Teacher Supe · ↓ older: From Routing Signals to Selective Review: Visual regrounding in MoE VL
Cite this page: LongHarness Bench: Stress-Testing Language Model Harnesses for Long-Context Reasoning: the #4 most recent of 292 cs.CL papers we have recorded (as of 2026-09-29). Source: Signals 4 (Signals API) — https://data.jiangzhang.ca/signals4/t/papers/longharness-bench-stress-testing-language-model-harnesses-for-long-context-reaso.html
Free to quote with attribution to “Signals 4 (Signals API)”. Machine-readable: papers.json
Related: More cs.CL papers · arXiv signals · All papers · Today in AI
Get 4 AI signals a day by email — free.
Subscribe free → See all plans →
Get 4 AI signals a day by email — free
All models · All repos · By company · Daily editions