Paper recorded by Signals 4 on 2026-08-31 in cs.CL. Abstract reproduced from arXiv; link to the original below.
Published 2026-08-31 on arXiv · recorded by Signals 4 on 2026-09-01
Category: cs.CL · 自然语言处理 · first seen 2026-09-01
Research planning is the decisive capability of AI scientists. Yet a research plan admits no verifiable answer, so reinforcement learning lacks the environment it requires: tasks paired with a critic. Rubrics extracted from scientific papers can supply the critic. Existing pipelines, however, draw the question and the criteria from the same content, so the reward can be earned by paraphrase. The r