Paper recorded by Signals 4 on 2026-09-09 in cs.AI. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-09 on arXiv · recorded by Signals 4 on 2026-09-10
Category: cs.AI · 人工智能 · first seen 2026-09-10
In high-stakes domains such as legal practice, a language-model answer is only useful to the extent that a reader can verify each claim against the source the system cites. Current grounded-generation pipelines score the answer as a whole, so a correct conclusion can rest on fabricated or loosely matched citations and still score well. Closing this gap requires both a system built for per-claim ve