Jailbreaking Large Vision-Language Models in Intelligent Transportation Systems

Jailbreaking Large Vision-Language Models in Intelligent Transportation Systems

Abstract

Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in multimodal reasoning and are increasingly deployed in real-world applications, including Intelligent Transportation Systems (ITS). However, these models remain highly vulnerable to jailbreak attacks that circumvent built-in safety mechanisms. This paper presents a systematic security analysis of LVLMs deployed in ITS under carefully crafted jailbreak attacks. We first construct a transportation-specific benchmark of harmful multimodal queries based on OpenAI’s prohibited content categories. We then propose a novel jailbreak attack that combines image typography manipulation with multi-turn prompting to conceal malicious intent while steering the model toward unsafe responses. To mitigate these attacks, we introduce a multi-layered response filtering defense that integrates rule-based filtering with a zero-shot classifier. Extensive experiments on both open-source and commercial LVLMs demonstrate the effectiveness of the proposed attack and defense. Attack success is evaluated using GPT-4-based toxicity assessment together with manual verification, and comparisons with existing jailbreak techniques highlight the significant security risks posed by image typography manipulation and multi-turn prompting in LVLM-enabled intelligent transportation applications.

Publication
2025 IEEE International Conference on Machine Learning and Applications (ICMLA)
Md Jueal Mia
Md Jueal Mia
Graduate Research Assistant

My research interests include Trustworthy AI, AI Security, Foundation Models, Large Language Models, Vision-Language Models, Agentic AI, Federated Learning, Privacy-Preserving Machine Learning, and Adversarial Machine Learning.