Signals 4 · free daily AI digest

When Does Predictor-Based RL Align with Human Perception? A Study of Subjective Rewards in Codec-Based Speech Language Models

Paper recorded by Signals 4 on 2026-08-31 in cs.CL. Abstract reproduced from arXiv; link to the original below.

Published 2026-08-31 on arXiv · recorded by Signals 4 on 2026-09-01

Category: cs.CL · 自然语言处理 · first seen 2026-09-01

Abstract

Codec-based text-to-speech (TTS) models make language-model post-training applicable to speech generation, but it remains unclear when learned perceptual predictors can serve as reinforcement learning rewards without losing alignment with human listeners. We study this question with Group Relative Policy Optimization (GRPO) using learned rewards for anime-like speaking style, naturalness, likabili

Read on arXiv →

#160 most recent of 186 cs.CL papers we have recorded · ↑ newer: Language-Statistical Analysis of Neural Audio Codec Tokens Across Arch · ↓ older: PrivBench: A Holistic and Modular Benchmarking Platform for Evaluating
Cite this page: When Does Predictor-Based RL Align with Human Perception? A Study of Subjective Rewards in Codec-Based Speech Language Models: the #160 most recent of 186 cs.CL papers we have recorded (as of 2026-08-31). Source: Signals 4 (Signals API) — https://data.jiangzhang.ca/signals4/t/papers/when-does-predictor-based-rl-align-with-human-perception-a-study-of-subjective-r.html
Free to quote with attribution to “Signals 4 (Signals API)”. Machine-readable: papers.json
Related: More cs.CL papers · arXiv signals · All papers · Today in AI
Get 4 AI signals a day by email — free.
Subscribe free → See all plans →
Get 4 AI signals a day by email — free
All models · All repos · By company · Daily editions