A Systematic Review of Trustworthiness, Hallucination, and Safety in Large Vision-Language Models
Subject Areas : AI and RoboticsDelaram Kiani 1 , Hanieh Naderi 2 *
1 - School of Intelligent Systems Engineering, College of Interdisciplinary Science and Technology, University of Tehran, Tehran, Iran
2 -
Keywords: Vision–Language Models, Reliability, Hallucination, Multimodal Safety, Visual Jailbreak.,
Abstract :
Large Vision–Language Models (LVLMs), despite their remarkable progress in multimodal tasks, continue to face two fundamental reliability challenges: the generation of hallucinatory content inconsistent with visual evidence, and vulnerability to adversarial attacks that bypass safety mechanisms. This systematic review, conducted in accordance with the PRISMA 2020 reporting framework, covers studies published from January 2022 to June 2026 and synthesizes 31 eligible studies. It investigates the technical roots of hallucination across multiple levels (object, attribute, relation, and narrative), visual jailbreak mechanisms, and mitigation strategies at both training and inference stages. The findings indicate that over-reliance on linguistic priors is a common underlying factor behind many hallucination errors and safety failures, and that purely text-based alignment is insufficient for multimodal systems. Furthermore, a significant gap exists between current evaluation capabilities and the real-world complexity of vulnerabilities, hindering comprehensive reliability assessment. This review provides an integrated taxonomy of vulnerabilities and outlines a research roadmap, offering a structured foundation for the development of more reliable vision–language models.
[1] T. Guan, F. Liu, X. Wu, R. Xian, Z. Li, X. Liu, X. Wang, L. Chen, F. Huang, Y. Yacoob, D. Manocha, and T. Zhou, “HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models,” Mar. 25, 2024, arXiv:2310.14566. doi: 10.48550/arXiv.2310.14566.
[2] Y. Li, Y. Du, K. Zhou, J. Wang, W. X. Zhao, and J.-R. Wen, “Evaluating Object Hallucination in Large Vision-Language Models,” Oct. 26, 2023, arXiv: arXiv:2305.10355. doi: 10.48550/arXiv.2305.10355.
[3] A. Seth, D. Manocha, and C. Agarwal, “Towards a Systematic Evaluation of Hallucinations in Large-Vision Language Models,” Mar. 13, 2025, arXiv: arXiv:2412.20622. doi: 10.48550/arXiv.2412.20622.
[4] Z. Liu, Y. Nie, Y. Tan, X. Yue, Q. Cui, C. Wang, X. Zhu, and B. Zheng, “Safety Alignment for Vision Language Models,” May 22, 2024, arXiv:2405.13581. doi: 10.48550/arXiv.2405.13581.
[5] W. Luo, S. Ma, X. Liu, X. Guo, and C. Xiao, “JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks,” Nov. 24, 2024, arXiv: arXiv:2404.03027. doi: 10.48550/arXiv.2404.03027.
[6] Z. Yang, J. Fan, A. Yan, E. Gao, X. Lin, T. Li, K. Mo, and C. Dong, “Distraction is All You Need for Multimodal Large Language Model Jailbreaking,” Jun. 17, 2025, arXiv:2502.10794. doi: 10.48550/arXiv.2502.10794.
[7] C. Jiang, H. Jia, W. Ye, M. Dong, H. Xu, M. Yan, J. Zhang, and S. Zhang, “Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models,” in Proceedings of the 32nd ACM International Conference on Multimedia, Melbourne, VIC, Australia: ACM, Oct. 2024, pp. 525–534. doi: 10.1145/3664647.3680576.
[8] Q. Cao, J. Cheng, X. Liang, and L. Lin, “VisDiaHalBench: A Visual Dialogue Benchmark For Diagnosing Hallucination in Large Vision-Language Models,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand: Association for Computational Linguistics, 2024, pp. 12161–12176. doi: 10.18653/v1/2024.acl-long.658.
[9] S. Schrodi, D. T. Hoffmann, M. Argus, V. Fischer, and T. Brox, “Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models,” Apr. 16, 2025, arXiv: arXiv:2404.07983. doi: 10.48550/arXiv.2404.07983.
[10] S. Yin, C. Fu, S. Zhao, T. Xu, H. Wang, D. Sui, Y. Shen, K. Li, X. Sun, and E. Chen, “Woodpecker: Hallucination Correction for Multimodal Large Language Models,” Science China Information Sciences, vol. 67, no. 12, p. 220105, Dec. 2024. doi: 10.1007/s11432-024-4251-x.
[11] V. Rawte, A. Mishra, A. Sheth, and A. Das, “Defining and quantifying visual hallucinations in vision-language models,” in Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025), 2025, pp. 501–510.
[12] W. Liang, Y. Zhang, Y. Kwon, S. Yeung, and J. Zou, “Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning,” Oct. 19, 2022, arXiv: arXiv:2203.02053. doi: 10.48550/arXiv.2203.02053.
[13] C. Yi, Y.-H. He, D.-C. Zhan, and H.-J. Ye, “Bridge the Modality and Capability Gaps in Vision-Language Model Selection,” May 18, 2025, arXiv: arXiv:2403.13797. doi: 10.48550/arXiv.2403.13797.
[14] X. Liu, Y. Zhu, J. Gu, Y. Lan, C. Yang, and Y. Qiao, “MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models,” Jun. 19, 2024, arXiv: arXiv:2311.17600. doi: 10.48550/arXiv.2311.17600.
[15] Z. Liu, Y. Nie, Y. Tan, J. Liu, X. Yue, Q. Cui, C. Wang, X. Zhu, and B. Zheng, “PSA-VLM: Enhancing Vision-Language Model Safety through Progressive Concept-Bottleneck-Driven Alignment,” 2024, arXiv:2411.11543. doi: 10.48550/arXiv.2411.11543.
[16] L. Huang, W. Yu, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qin, and T. Liu, “A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions,” 2023, arXiv:2311.05232. doi: 10.48550/arXiv.2311.05232.
[17] Z. Bai, P. Wang, T. Xiao, T. He, Z. Han, Z. Zhang, and M. Z. Shou, “Hallucination of Multimodal Large Language Models: A Survey,” Apr. 01, 2025, arXiv:2404.18930. doi: 10.48550/arXiv.2404.18930.
[18] Z. Yang, X. Luo, D. Han, Y. Xu, and D. Li, “Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key,” 2025, arXiv. doi: 10.48550/ARXIV.2501.09695.
[19] P. Kaul, Z. Li, H. Yang, Y. Dukler, A. Swaminathan, C. J. Taylor, and S. Soatto, “THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models,” Apr. 03, 2025, arXiv:2405.05256. doi: 10.48550/arXiv.2405.05256.
[20] X. Fan, X. Wang, H. Wei, X. Zhang, and D. Zhao, “Toward a Stable, Fair, and Comprehensive Evaluation of Object Hallucination in Large Vision-Language Models,” in Advances in Neural Information Processing Systems 37, Vancouver, BC, Canada: Neural Information Processing Systems Foundation, Inc. (NeurIPS), 2024, pp. 111406–111431. doi: 10.52202/079017-3538.
[21] M. Ye-Bin, N. Hyeon-Woo, W. Choi, and T.-H. Oh, “BEAF: Observing BEfore-AFter Changes to Evaluate Hallucination in Vision-language Models,” Jul. 18, 2024, arXiv: arXiv:2407.13442. doi: 10.48550/arXiv.2407.13442.
[22] M. Ye, X. Rong, W. Huang, B. Du, N. Yu, and D. Tao, “A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations,” Feb. 14, 2025, arXiv: arXiv:2502.14881. doi: 10.48550/arXiv.2502.14881.
[23] E. Shayegani, Y. Dong, and N. Abu-Ghazaleh, “Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models,” Oct. 10, 2023, arXiv: arXiv:2307.14539. doi: 10.48550/arXiv.2307.14539.
[24] W. Hu, S. Gu, Y. Wang, and R. Hong, “VideoJail: Exploiting Video-Modality Vulnerabilities for Jailbreak Attacks on Multimodal Large Language Models,” presented at the ICLR 2025 Workshop on Building Trust in Language Models and Applications, Mar. 2025. Accessed: Feb. 03, 2026. [Online]. Available: https://openreview.net/forum?id=fSAIDcPduZ
[25] H. Zhong, Q. Teng, B. Zheng, G. Chen, Y. Tan, Z. Liu, J. Liu, W. Su, X. Zhu, B. Zheng, and K. Zhang, “Towards Visualization-of-Thought Jailbreak Attack against Large Visual Language Models,” presented at the Thirty-ninth Annual Conference on Neural Information Processing Systems, Oct. 2025. Accessed: Feb. 03, 2026. [Online]. Available: https://openreview.net/forum?id=5P5YgohyBZ.
[26] S. S. Ghosal, S. Chakraborty, V. Singh, T. Guan, M. Wang, A. Velasquez, A. Beirami, F. Huang, D. Manocha, and A. S. Bedi, “Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment,” Jun. 14, 2025, arXiv:2411.18688. doi: 10.48550/arXiv.2411.18688.
[27] B. Chen, X. Lyu, L. Gao, J. Song, and H. T. Shen, “SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore Mechanism,” Dec. 03, 2025, arXiv: arXiv:2507.01513. doi: 10.48550/arXiv.2507.01513.
[28] W. You, Z. Li, H. Guo, X. Liu, W. Wang, X. Dong, and G. Wang, “MIRAGE: Multimodal Immersive Reasoning and Guided Exploration for Red-Team Jailbreak Attacks,” Mar. 24, 2025, arXiv:2503.19134. doi: 10.48550/arXiv.2503.19134.
[29] L. Zhu, D. Ji, T. Chen, P. Xu, J. Ye, and J. Liu, “IBD: Alleviating Hallucinations in Large Vision-Language Models via Image-Biased Decoding,” 2024, arXiv:2402.18476. doi: 10.48550/arXiv.2402.18476.
[30] Y. Shu, H. Lin, Y. Liu, Y. Zhang, G. Zeng, Y. Li, Y. Zhou, S.-N. Lim, H. Yang, and N. Sebe, “When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding,” Oct. 07, 2025, arXiv:2506.05551. doi: 10.48550/arXiv.2506.05551.
[31] X. Liu, M. Luo, A. Chatterjee, H. Wei, C. Baral, and Y. Yang, “Investigating VLM Hallucination from a Cognitive Psychology Perspective: A First Step Toward Interpretation with Intriguing Observations,” arXiv preprint arXiv:2507.03123, 2025.
[32] E. Shayegani, Y. Dong, and N. Abu-Ghazaleh, “Jailbreak in Pieces: Compositional Adversarial Attacks on Multi-Modal Language Models,” Oct. 10, 2023, arXiv:2307.14539. doi: 10.48550/arXiv.2307.14539.
[33] L. He, Z. Chen, Z. Shi, T. Yu, J. Shao, and L. Sheng, “Systematic Reward Gap Optimization for Mitigating VLM Hallucinations,” Nov. 24, 2025, arXiv:2411.17265. doi: 10.48550/arXiv.2411.17265.
[34] X. Lyu, B. Chen, L. Gao, J. Song, and H. T. Shen, “Alleviating Hallucinations in Large Vision-Language Models through Hallucination-Induced Optimization,” Apr. 01, 2025, arXiv: arXiv:2405.15356. doi: 10.48550/arXiv.2405.15356.
[35] Y. Shu, H. Lin, Y. Liu, Y. Zhang, G. Zeng, Y. Li, Y. Zhou, S.-N. Lim, H. Yang, and N. Sebe, “TextHalu-Bench and Semantic Hallucination Mitigation for Scene Text Understanding in Large Multimodal Models,” 2025, arXiv:2506.05551. doi: 10.48550/arXiv.2506.05551.
[36] Z. Sun, S. Shen, S. Cao, H. Liu, C. Li, Y. Shen, C. Gan, L.-Y. Gui, Y.-X. Wang, Y. Yang, K. Keutzer, and T. Darrell, “Aligning Large Multimodal Models with Factually Augmented RLHF,” Sep. 25, 2023, arXiv:2309.14525. doi: 10.48550/arXiv.2309.14525.
[37] Y. Xie, G. Li, X. Xu, and M.-Y. Kan, “V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization,” Nov. 05, 2024, arXiv: arXiv:2411.02712. doi: 10.48550/arXiv.2411.02712.
[38] Z. Yang, X. Luo, D. Han, Y. Xu, and D. Li, “Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key,” Mar. 03, 2025, arXiv: arXiv:2501.09695. doi: 10.48550/arXiv.2501.09695.
[39] N. Jiang, A. Kachinthaya, S. Petryk, and Y. Gandelsman, “Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations,” Feb. 10, 2025, arXiv: arXiv:2410.02762. doi: 10.48550/arXiv.2410.02762.
[40] X. Chen, Z. Ma, X. Zhang, S. Xu, S. Qian, J. Yang, D. F. Fouhey, and J. Chai, “Multi-Object Hallucination in Vision-Language Models,” Oct. 31, 2024, arXiv:2407.06192. doi: 10.48550/arXiv.2407.06192.