Signals 4 · free daily AI digest

$λ$-Controlled GRPO: Turning Flow-Matching Ratio Instability into a Budgeted Resource

Paper recorded by Signals 4 on 2026-09-18 in cs.LG. Abstract reproduced from arXiv; link to the original below.

Published 2026-09-18 on arXiv · recorded by Signals 4 on 2026-09-21

Category: cs.LG · 机器学习 · first seen 2026-09-21

Abstract

Reinforcement learning is increasingly used to align image generators with reward signals, and Flow-GRPO recently extended this paradigm to flow-matching models by treating the denoising sampler as a stochastic policy that can be optimized from reward feedback. Training in this setting is unstable in a way specific to multi-step denoising: the policy update changes systematically across denoising

Read on arXiv →

#6 most recent of 235 cs.LG papers we have recorded · ↑ newer: Available Guardrails: Certifying Selective Prediction across ML System · ↓ older: COMPLEX: A Closed-Form Certified Embedding of Multiparameter Persisten
Cite this page: $λ$-Controlled GRPO: Turning Flow-Matching Ratio Instability into a Budgeted Resource: the #6 most recent of 235 cs.LG papers we have recorded (as of 2026-09-18). Source: Signals 4 (Signals API) — https://data.jiangzhang.ca/signals4/t/papers/controlled-grpo-turning-flow-matching-ratio-instability-into-a-budgeted-resource.html
Free to quote with attribution to “Signals 4 (Signals API)”. Machine-readable: papers.json
Related: More cs.LG papers · arXiv signals · All papers · Today in AI
Get 4 AI signals a day by email — free.
Subscribe free → See all plans →
Get 4 AI signals a day by email — free
All models · All repos · By company · Daily editions