Paper recorded by Signals 4 on 2026-09-11 in cs.LG. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-11 on arXiv · recorded by Signals 4 on 2026-09-14
Category: cs.LG · 机器学习 · first seen 2026-09-14
Significant progress has been made in safeguarding deep reinforcement learning (DRL) policies against input perturbations. Developing robust DRL involves three main stages: algorithm design, implementation, and evaluation. In this work, we identify and address a key limitation at each stage. First, we introduce Adversarial Importance Sampling (Advis), a method that uses importance sampling over tr