Paper recorded by Signals 4 on 2026-09-03 in cs.AI. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-03 on arXiv · recorded by Signals 4 on 2026-09-04
Category: cs.AI · 人工智能 · first seen 2026-09-04
Repository-level software engineering benchmarks have significantly advanced the evaluation of coding agents, but existing benchmarks primarily measure whether generated patches pass functional tests and overlook review-derived acceptance constraints (review constraints) that often influence whether a patch is acceptable in real-world software development. We introduce SWE-Gate, a repository-level