When Does Predictor-Based RL Align with Human Perception? A Study of Subjective Rewards in Codec-Based Speech Language Models
Paper recorded by Signals 4 on 2026-08-31 in cs.CL. Abstract reproduced from arXiv; link to the original below.
Published 2026-08-31 on arXiv · recorded by Signals 4 on 2026-09-01
Category: cs.CL · 自然语言处理 · first seen 2026-09-01
Abstract
Codec-based text-to-speech (TTS) models make language-model post-training applicable to speech generation, but it remains unclear when learned perceptual predictors can serve as reinforcement learning rewards without losing alignment with human listeners. We study this question with Group Relative Policy Optimization (GRPO) using learned rewards for anime-like speaking style, naturalness, likabili
Read on arXiv →
Cite this page: When Does Predictor-Based RL Align with Human Perception? A Study of Subjective Rewards in Codec-Based Speech Language Models: the #160 most recent of 186 cs.CL papers we have recorded (as of 2026-08-31). Source: Signals 4 (Signals API) — https://data.jiangzhang.ca/signals4/t/papers/when-does-predictor-based-rl-align-with-human-perception-a-study-of-subjective-r.html
Free to quote with attribution to “Signals 4 (Signals API)”. Machine-readable:
papers.json
Get 4 AI signals a day by email — free.
Get 4 AI signals a day by email — free