Signals 4 · free daily AI digest

CASCADE Against Jailbreaks: Combination Across Stages with Controlled Attack-Defense Evaluation

Paper recorded by Signals 4 on 2026-09-18 in cs.CL. Abstract reproduced from arXiv; link to the original below.

Published 2026-09-18 on arXiv · recorded by Signals 4 on 2026-09-21

Category: cs.CL · 自然语言处理 · first seen 2026-09-21

Abstract

Defenses against jailbreak attacks on Large Language Models (LLMs) operate at different pipeline stages, such as input modification or output guard, but it remains unclear which defenses to deploy at each stage and how to combine them. Prior empirical studies, fragmented by inconsistent attack-success-rate definitions and experimental settings, have evaluated defenses largely in isolation. Here we

Read on arXiv →

#9 most recent of 200 cs.CL papers we have recorded · ↑ newer: RheoSampling: Resolving the One-Hot Dilemma in Stochastic Dynamic-Tree · ↓ older: Per-Aetiology Contrastive Severity Embeddings with Phonological Pseudo
Cite this page: CASCADE Against Jailbreaks: Combination Across Stages with Controlled Attack-Defense Evaluation: the #9 most recent of 200 cs.CL papers we have recorded (as of 2026-09-18). Source: Signals 4 (Signals API) — https://data.jiangzhang.ca/signals4/t/papers/cascade-against-jailbreaks-combination-across-stages-with-controlled-attack-defe.html
Free to quote with attribution to “Signals 4 (Signals API)”. Machine-readable: papers.json
Related: More cs.CL papers · arXiv signals · All papers · Today in AI
Get 4 AI signals a day by email — free.
Subscribe free → See all plans →
Get 4 AI signals a day by email — free
All models · All repos · By company · Daily editions