Signals 4 · free daily AI digest

Adaptive Doubly Robust Off-Policy Evaluation for Ranking Policies under Diverse User Behavior

Paper recorded by Signals 4 on 2026-08-30 in cs.LG. Abstract reproduced from arXiv; link to the original below.

Published 2026-08-30 on arXiv · recorded by Signals 4 on 2026-09-01

Category: cs.LG · 机器学习 · first seen 2026-09-01

Abstract

Off-policy evaluation (OPE) of ranking policies is challenging be- cause selecting and ordering multiple items from a candidate set makes the number of possible rankings grow combinatorially with the number of candidates and the ranking length. Consequently, Inverse Propensity Scoring (IPS), whose importance weight is the full-ranking probability ratio under the evaluation and logging policies, ca

Read on arXiv →

#190 most recent of 215 cs.LG papers we have recorded · ↑ newer: $\mathcal{N}_0$-Foundation: Towards the Age of Tactile Intelligence · ↓ older: On the Resilience of Text-to-Video Diffusion Models to Hardware Faults
Cite this page: Adaptive Doubly Robust Off-Policy Evaluation for Ranking Policies under Diverse User Behavior: the #190 most recent of 215 cs.LG papers we have recorded (as of 2026-08-30). Source: Signals 4 (Signals API) — https://data.jiangzhang.ca/signals4/t/papers/adaptive-doubly-robust-off-policy-evaluation-for-ranking-policies-under-diverse-.html
Free to quote with attribution to “Signals 4 (Signals API)”. Machine-readable: papers.json
Related: More cs.LG papers · arXiv signals · All papers · Today in AI
Get 4 AI signals a day by email — free.
Subscribe free → See all plans →
Get 4 AI signals a day by email — free
All models · All repos · By company · Daily editions