SAVeS: Steering Safety Judgments in Vision-Language Models via Semantic Cues
Carlos Hinojosa, Clemens Grange, Bernard Ghanem
This paper studies how vision-language models (VLMs) make safety decisions and finds that they can be easily manipulated by changing visual or textual cues without altering the actual scene content. The researchers created a benchmark called SAVeS to test these safety decisions and discovered that VLMs rely on learned associations between images and text rather than truly understanding the visual content, which reveals a potential security vulnerability in these AI systems.
vision-language modelssafetyadversarial robustnessinterpretability