Return
SafeSteer: Adaptive Subspace Steering for Efficient Jailbreak Defense in Vision Language Models
S
X
S
L
H
C
E
L
B
Y
J
DOI:10.1109/tifs.2026.3714149.png)
Abstract
En 中文
As the capabilities of Vision Language Models (VLMs) continue to improve, they are increasingly targeted by jailbreak attacks. Existing defense methods face two major limitations: (1) they struggle to ensure safety without compromising the model’s utility; and (2) many defense mechanisms significantly reduce the model’s generation efficiency. To address these challenges, we propose SafeSteer, a lightweight inference-time steering framework that effectively defends against diverse jailbreak attacks without modifying model weights. At the core of SafeSteer is the innovative use of singular value decomposition (SVD) to purify a low-dimensional “safety subspace” from noisy activation differences. By projecting the raw steering vector into this subspace, SafeSteer isolates the core safety signal from noise, adaptively removing harmful influences while preserving the model’s ability to handle benign inputs. SafeSteer avoids iterative response generation and introduces only limited overhead compared with other single-pass activation-steering defenses. Extensive experiments show that SafeSteer reduces the attack success rate by over 60% while maintaining the model’s utility on benign tasks, without introducing significant inference latency. These results demonstrate that robust and practical jailbreak defense can be achieved through simple, efficient inference-time control.
Keywords:
Vision language models
jailbreak defense
activation steering
inference-time intervention
Journal
IF:
8
Papers:
5.2K
Citations:
2.3W
