مروری نظام مند بر قابلیت اعتماد، توهم و ایمنی در مدلهای بزرگ بینایی-زبان
محورهای موضوعی : هوش مصنوعی و رباتیکدلارام کیانی 1 , حانیه نادری 2 *
1 - دانشکده مهندسی سامانههای هوشمند، دانشکدگان علوم و فناوریهای میانرشتهای، دانشگاه تهران، تهران، ایران
2 - دانشکده مهندسی سامانههای هوشمند، دانشکدگان علوم و فناوریهای میانرشتهای، دانشگاه تهران، تهران، ایران
کلید واژه: مدلهای بینایی-زبان, قابلیت اعتماد, توهم, ایمنی چندوجهی, جیلبریک بصری,
چکیده مقاله :
مدلهای بزرگ بینایی-زبان با وجود پیشرفتهای چشمگیر در وظایف چندوجهی، همچنان با دو چالش بنیادین قابلیت اعتماد روبهرو هستند: تولید محتوای توهمی ناسازگار با شواهد بصری، و آسیبپذیری در برابر حملات خصمانه که سازوکارهای ایمنی را دور میزنند. این مرور نظاممند، با پیروی از چارچوب گزارشدهی پریزما ۲۰۲۰، مطالعات منتشرشده از ژانویه ۲۰۲۲ تا ژوئن ۲۰۲۶ را پوشش میدهد و ۳۱ مطالعه واجد شرایط را ترکیب و تحلیل میکند. این مطالعه ریشههای فنی توهم در سطوح مختلف شامل شیء، ویژگی، رابطه و روایت، سازوکارهای گریز امنیتی بصری و راهبردهای کاهش در زمان آموزش و استنتاج را بررسی میکند. یافتههای این مطالعه نشان میدهد که اتکای بیش از حد به پیشینهای زبانی، ریشه مشترک بسیاری از خطاهای توهمی و شکستهای ایمنی است و همترازی صرفاً متنی برای سیستمهای چندوجهی کافی نیست. همچنین، شکاف معناداری میان توانمندیهای ارزیابی فعلی و پیچیدگی واقعی آسیبپذیریها وجود دارد که مانع سنجش جامع قابلیت اعتماد میشود. این مرور با ارائه ردهبندی یکپارچه از آسیبپذیریها و نقشه راه پژوهشی، چارچوبی برای توسعه مدلهای قابل اعتمادتر فراهم میآورد.
Large Vision–Language Models (LVLMs), despite their remarkable progress in multimodal tasks, continue to face two fundamental reliability challenges: the generation of hallucinatory content inconsistent with visual evidence, and vulnerability to adversarial attacks that bypass safety mechanisms. This systematic review, conducted in accordance with the PRISMA 2020 reporting framework, covers studies published from January 2022 to June 2026 and synthesizes 31 eligible studies. It investigates the technical roots of hallucination across multiple levels (object, attribute, relation, and narrative), visual jailbreak mechanisms, and mitigation strategies at both training and inference stages. The findings indicate that over-reliance on linguistic priors is a common underlying factor behind many hallucination errors and safety failures, and that purely text-based alignment is insufficient for multimodal systems. Furthermore, a significant gap exists between current evaluation capabilities and the real-world complexity of vulnerabilities, hindering comprehensive reliability assessment. This review provides an integrated taxonomy of vulnerabilities and outlines a research roadmap, offering a structured foundation for the development of more reliable vision–language models.
[1] T. Guan, F. Liu, X. Wu, R. Xian, Z. Li, X. Liu, X. Wang, L. Chen, F. Huang, Y. Yacoob, D. Manocha, and T. Zhou, “HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models,” Mar. 25, 2024, arXiv:2310.14566. doi: 10.48550/arXiv.2310.14566.
[2] Y. Li, Y. Du, K. Zhou, J. Wang, W. X. Zhao, and J.-R. Wen, “Evaluating Object Hallucination in Large Vision-Language Models,” Oct. 26, 2023, arXiv: arXiv:2305.10355. doi: 10.48550/arXiv.2305.10355.
[3] A. Seth, D. Manocha, and C. Agarwal, “Towards a Systematic Evaluation of Hallucinations in Large-Vision Language Models,” Mar. 13, 2025, arXiv: arXiv:2412.20622. doi: 10.48550/arXiv.2412.20622.
[4] Z. Liu, Y. Nie, Y. Tan, X. Yue, Q. Cui, C. Wang, X. Zhu, and B. Zheng, “Safety Alignment for Vision Language Models,” May 22, 2024, arXiv:2405.13581. doi: 10.48550/arXiv.2405.13581.
[5] W. Luo, S. Ma, X. Liu, X. Guo, and C. Xiao, “JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks,” Nov. 24, 2024, arXiv: arXiv:2404.03027. doi: 10.48550/arXiv.2404.03027.
[6] Z. Yang, J. Fan, A. Yan, E. Gao, X. Lin, T. Li, K. Mo, and C. Dong, “Distraction is All You Need for Multimodal Large Language Model Jailbreaking,” Jun. 17, 2025, arXiv:2502.10794. doi: 10.48550/arXiv.2502.10794.
[7] C. Jiang, H. Jia, W. Ye, M. Dong, H. Xu, M. Yan, J. Zhang, and S. Zhang, “Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models,” in Proceedings of the 32nd ACM International Conference on Multimedia, Melbourne, VIC, Australia: ACM, Oct. 2024, pp. 525–534. doi: 10.1145/3664647.3680576.
[8] Q. Cao, J. Cheng, X. Liang, and L. Lin, “VisDiaHalBench: A Visual Dialogue Benchmark For Diagnosing Hallucination in Large Vision-Language Models,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand: Association for Computational Linguistics, 2024, pp. 12161–12176. doi: 10.18653/v1/2024.acl-long.658.
[9] S. Schrodi, D. T. Hoffmann, M. Argus, V. Fischer, and T. Brox, “Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models,” Apr. 16, 2025, arXiv: arXiv:2404.07983. doi: 10.48550/arXiv.2404.07983.
[10] S. Yin, C. Fu, S. Zhao, T. Xu, H. Wang, D. Sui, Y. Shen, K. Li, X. Sun, and E. Chen, “Woodpecker: Hallucination Correction for Multimodal Large Language Models,” Science China Information Sciences, vol. 67, no. 12, p. 220105, Dec. 2024. doi: 10.1007/s11432-024-4251-x.
[11] V. Rawte, A. Mishra, A. Sheth, and A. Das, “Defining and quantifying visual hallucinations in vision-language models,” in Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025), 2025, pp. 501–510.
[12] W. Liang, Y. Zhang, Y. Kwon, S. Yeung, and J. Zou, “Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning,” Oct. 19, 2022, arXiv: arXiv:2203.02053. doi: 10.48550/arXiv.2203.02053.
[13] C. Yi, Y.-H. He, D.-C. Zhan, and H.-J. Ye, “Bridge the Modality and Capability Gaps in Vision-Language Model Selection,” May 18, 2025, arXiv: arXiv:2403.13797. doi: 10.48550/arXiv.2403.13797.
[14] X. Liu, Y. Zhu, J. Gu, Y. Lan, C. Yang, and Y. Qiao, “MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models,” Jun. 19, 2024, arXiv: arXiv:2311.17600. doi: 10.48550/arXiv.2311.17600.
[15] Z. Liu, Y. Nie, Y. Tan, J. Liu, X. Yue, Q. Cui, C. Wang, X. Zhu, and B. Zheng, “PSA-VLM: Enhancing Vision-Language Model Safety through Progressive Concept-Bottleneck-Driven Alignment,” 2024, arXiv:2411.11543. doi: 10.48550/arXiv.2411.11543.
[16] L. Huang, W. Yu, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qin, and T. Liu, “A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions,” 2023, arXiv:2311.05232. doi: 10.48550/arXiv.2311.05232.
[17] Z. Bai, P. Wang, T. Xiao, T. He, Z. Han, Z. Zhang, and M. Z. Shou, “Hallucination of Multimodal Large Language Models: A Survey,” Apr. 01, 2025, arXiv:2404.18930. doi: 10.48550/arXiv.2404.18930.
[18] Z. Yang, X. Luo, D. Han, Y. Xu, and D. Li, “Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key,” 2025, arXiv. doi: 10.48550/ARXIV.2501.09695.
[19] P. Kaul, Z. Li, H. Yang, Y. Dukler, A. Swaminathan, C. J. Taylor, and S. Soatto, “THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models,” Apr. 03, 2025, arXiv:2405.05256. doi: 10.48550/arXiv.2405.05256.
[20] X. Fan, X. Wang, H. Wei, X. Zhang, and D. Zhao, “Toward a Stable, Fair, and Comprehensive Evaluation of Object Hallucination in Large Vision-Language Models,” in Advances in Neural Information Processing Systems 37, Vancouver, BC, Canada: Neural Information Processing Systems Foundation, Inc. (NeurIPS), 2024, pp. 111406–111431. doi: 10.52202/079017-3538.
[21] M. Ye-Bin, N. Hyeon-Woo, W. Choi, and T.-H. Oh, “BEAF: Observing BEfore-AFter Changes to Evaluate Hallucination in Vision-language Models,” Jul. 18, 2024, arXiv: arXiv:2407.13442. doi: 10.48550/arXiv.2407.13442.
[22] M. Ye, X. Rong, W. Huang, B. Du, N. Yu, and D. Tao, “A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations,” Feb. 14, 2025, arXiv: arXiv:2502.14881. doi: 10.48550/arXiv.2502.14881.
[23] E. Shayegani, Y. Dong, and N. Abu-Ghazaleh, “Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models,” Oct. 10, 2023, arXiv: arXiv:2307.14539. doi: 10.48550/arXiv.2307.14539.
[24] W. Hu, S. Gu, Y. Wang, and R. Hong, “VideoJail: Exploiting Video-Modality Vulnerabilities for Jailbreak Attacks on Multimodal Large Language Models,” presented at the ICLR 2025 Workshop on Building Trust in Language Models and Applications, Mar. 2025. Accessed: Feb. 03, 2026. [Online]. Available: https://openreview.net/forum?id=fSAIDcPduZ
[25] H. Zhong, Q. Teng, B. Zheng, G. Chen, Y. Tan, Z. Liu, J. Liu, W. Su, X. Zhu, B. Zheng, and K. Zhang, “Towards Visualization-of-Thought Jailbreak Attack against Large Visual Language Models,” presented at the Thirty-ninth Annual Conference on Neural Information Processing Systems, Oct. 2025. Accessed: Feb. 03, 2026. [Online]. Available: https://openreview.net/forum?id=5P5YgohyBZ.
[26] S. S. Ghosal, S. Chakraborty, V. Singh, T. Guan, M. Wang, A. Velasquez, A. Beirami, F. Huang, D. Manocha, and A. S. Bedi, “Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment,” Jun. 14, 2025, arXiv:2411.18688. doi: 10.48550/arXiv.2411.18688.
[27] B. Chen, X. Lyu, L. Gao, J. Song, and H. T. Shen, “SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore Mechanism,” Dec. 03, 2025, arXiv: arXiv:2507.01513. doi: 10.48550/arXiv.2507.01513.
[28] W. You, Z. Li, H. Guo, X. Liu, W. Wang, X. Dong, and G. Wang, “MIRAGE: Multimodal Immersive Reasoning and Guided Exploration for Red-Team Jailbreak Attacks,” Mar. 24, 2025, arXiv:2503.19134. doi: 10.48550/arXiv.2503.19134.
[29] L. Zhu, D. Ji, T. Chen, P. Xu, J. Ye, and J. Liu, “IBD: Alleviating Hallucinations in Large Vision-Language Models via Image-Biased Decoding,” 2024, arXiv:2402.18476. doi: 10.48550/arXiv.2402.18476.
[30] Y. Shu, H. Lin, Y. Liu, Y. Zhang, G. Zeng, Y. Li, Y. Zhou, S.-N. Lim, H. Yang, and N. Sebe, “When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding,” Oct. 07, 2025, arXiv:2506.05551. doi: 10.48550/arXiv.2506.05551.
[31] X. Liu, M. Luo, A. Chatterjee, H. Wei, C. Baral, and Y. Yang, “Investigating VLM Hallucination from a Cognitive Psychology Perspective: A First Step Toward Interpretation with Intriguing Observations,” arXiv preprint arXiv:2507.03123, 2025.
[32] E. Shayegani, Y. Dong, and N. Abu-Ghazaleh, “Jailbreak in Pieces: Compositional Adversarial Attacks on Multi-Modal Language Models,” Oct. 10, 2023, arXiv:2307.14539. doi: 10.48550/arXiv.2307.14539.
[33] L. He, Z. Chen, Z. Shi, T. Yu, J. Shao, and L. Sheng, “Systematic Reward Gap Optimization for Mitigating VLM Hallucinations,” Nov. 24, 2025, arXiv:2411.17265. doi: 10.48550/arXiv.2411.17265.
[34] X. Lyu, B. Chen, L. Gao, J. Song, and H. T. Shen, “Alleviating Hallucinations in Large Vision-Language Models through Hallucination-Induced Optimization,” Apr. 01, 2025, arXiv: arXiv:2405.15356. doi: 10.48550/arXiv.2405.15356.
[35] Y. Shu, H. Lin, Y. Liu, Y. Zhang, G. Zeng, Y. Li, Y. Zhou, S.-N. Lim, H. Yang, and N. Sebe, “TextHalu-Bench and Semantic Hallucination Mitigation for Scene Text Understanding in Large Multimodal Models,” 2025, arXiv:2506.05551. doi: 10.48550/arXiv.2506.05551.
[36] Z. Sun, S. Shen, S. Cao, H. Liu, C. Li, Y. Shen, C. Gan, L.-Y. Gui, Y.-X. Wang, Y. Yang, K. Keutzer, and T. Darrell, “Aligning Large Multimodal Models with Factually Augmented RLHF,” Sep. 25, 2023, arXiv:2309.14525. doi: 10.48550/arXiv.2309.14525.
[37] Y. Xie, G. Li, X. Xu, and M.-Y. Kan, “V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization,” Nov. 05, 2024, arXiv: arXiv:2411.02712. doi: 10.48550/arXiv.2411.02712.
[38] Z. Yang, X. Luo, D. Han, Y. Xu, and D. Li, “Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key,” Mar. 03, 2025, arXiv: arXiv:2501.09695. doi: 10.48550/arXiv.2501.09695.
[39] N. Jiang, A. Kachinthaya, S. Petryk, and Y. Gandelsman, “Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations,” Feb. 10, 2025, arXiv: arXiv:2410.02762. doi: 10.48550/arXiv.2410.02762.
[40] X. Chen, Z. Ma, X. Zhang, S. Xu, S. Qian, J. Yang, D. F. Fouhey, and J. Chai, “Multi-Object Hallucination in Vision-Language Models,” Oct. 31, 2024, arXiv:2407.06192. doi: 10.48550/arXiv.2407.06192.