Adaptive Doubly Robust Off-Policy Evaluation for Ranking Policies under Diverse User Behavior
Paper recorded by Signals 4 on 2026-08-30 in cs.LG. Abstract reproduced from arXiv; link to the original below.
Published 2026-08-30 on arXiv · recorded by Signals 4 on 2026-09-01
Category: cs.LG · 机器学习 · first seen 2026-09-01
Abstract
Off-policy evaluation (OPE) of ranking policies is challenging be- cause selecting and ordering multiple items from a candidate set makes the number of possible rankings grow combinatorially with the number of candidates and the ranking length. Consequently, Inverse Propensity Scoring (IPS), whose importance weight is the full-ranking probability ratio under the evaluation and logging policies, ca
Read on arXiv →
Cite this page: Adaptive Doubly Robust Off-Policy Evaluation for Ranking Policies under Diverse User Behavior: the #190 most recent of 215 cs.LG papers we have recorded (as of 2026-08-30). Source: Signals 4 (Signals API) — https://data.jiangzhang.ca/signals4/t/papers/adaptive-doubly-robust-off-policy-evaluation-for-ranking-policies-under-diverse-.html
Free to quote with attribution to “Signals 4 (Signals API)”. Machine-readable:
papers.json
Get 4 AI signals a day by email — free.
Get 4 AI signals a day by email — free