Signals 4 · free daily AI digest

SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents

Paper recorded by Signals 4 on 2026-09-03 in cs.AI. Abstract reproduced from arXiv; link to the original below.

Published 2026-09-03 on arXiv · recorded by Signals 4 on 2026-09-04

Category: cs.AI · 人工智能 · first seen 2026-09-04

Abstract

Repository-level software engineering benchmarks have significantly advanced the evaluation of coding agents, but existing benchmarks primarily measure whether generated patches pass functional tests and overlook review-derived acceptance constraints (review constraints) that often influence whether a patch is acceptable in real-world software development. We introduce SWE-Gate, a repository-level

Read on arXiv →

#190 most recent of 300 cs.AI papers we have recorded · ↑ newer: A Case Study on Emergent Cheating and Whistleblowing in Autonomous Res · ↓ older: From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for
Cite this page: SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents: the #190 most recent of 300 cs.AI papers we have recorded (as of 2026-09-03). Source: Signals 4 (Signals API) — https://data.jiangzhang.ca/signals4/t/papers/swe-gate-passing-functional-tests-is-not-enough-for-software-engineering-agents.html
Free to quote with attribution to “Signals 4 (Signals API)”. Machine-readable: papers.json
Related: More cs.AI papers · arXiv signals · All papers · Today in AI
Get 4 AI signals a day by email — free.
Subscribe free → See all plans →
Get 4 AI signals a day by email — free
All models · All repos · By company · Daily editions