Paper recorded by Signals 4 on 2026-09-01 in cs.LG. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-01 on arXiv · recorded by Signals 4 on 2026-09-02
Category: cs.LG · 机器学习 · first seen 2026-09-02
LLM-based evaluators of natural language generation (NLG) quality are widely deployed as scoring tools and as automated training signals, yet the internal procedure by which they assign a rating remains poorly understood. We investigate this procedure mechanistically through an eight-attack perturbation taxonomy across the Readability and Adequacy dimensions of NLG quality, a generation pipeline t