Sycophantic Agreement Transfers with Neutral Data via Contrastive Preference Optimization
Paper recorded by Signals 4 on 2026-08-31 in cs.LG. Abstract reproduced from arXiv; link to the original below.
Published 2026-08-31 on arXiv · recorded by Signals 4 on 2026-09-01
Category: cs.LG · 机器学习 · first seen 2026-09-01
Abstract
Sycophantic agreement refers to a behavior in which language models excessively affirm the user, often at the cost of factual accuracy. Although sycophantic agreement is a well-known failure of model alignment, there is limited understanding of how it emerges from model training. In this work, we demonstrate that sycophantic agreement can emerge as an unintended consequence of widely used contrast
Read on arXiv →
Cite this page: Sycophantic Agreement Transfers with Neutral Data via Contrastive Preference Optimization: the #176 most recent of 215 cs.LG papers we have recorded (as of 2026-08-31). Source: Signals 4 (Signals API) — https://data.jiangzhang.ca/signals4/t/papers/sycophantic-agreement-transfers-with-neutral-data-via-contrastive-preference-opt.html
Free to quote with attribution to “Signals 4 (Signals API)”. Machine-readable:
papers.json
Get 4 AI signals a day by email — free.
Get 4 AI signals a day by email — free