A Systematic Review of Trustworthiness, Hallucination, and Safety in Large Vision-Language Models
Hanieh Naderi
1
(
)
Delaram Kiani
2
(
School of Intelligent Systems Engineering, College of Interdisciplinary Science and Technology, University of Tehran, Tehran, Iran
)
Keywords: Vision–Language Models, Reliability, Hallucination, Multimodal Safety, Visual Jailbreak.,
Abstract :
Large Vision–Language Models (LVLMs), despite their remarkable progress in multimodal tasks, continue to face two fundamental reliability challenges: the generation of hallucinatory content inconsistent with visual evidence, and vulnerability to adversarial attacks that bypass safety mechanisms. This systematic review, conducted in accordance with the PRISMA 2020 reporting framework and based on 31 selected studies from 2023 to 2025, investigates the technical roots of hallucination across multiple levels (object, attribute, relation, and narrative), visual jailbreak mechanisms, and mitigation strategies at both training and inference stages.
The findings indicate that over-reliance on linguistic priors is a common underlying factor behind many hallucination errors and safety failures, and that purely text-based alignment is insufficient for multimodal systems. Furthermore, a significant gap exists between current evaluation capabilities and the real-world complexity of vulnerabilities, hindering comprehensive reliability assessment.
This review provides an integrated taxonomy of vulnerabilities and outlines a research roadmap, offering a structured foundation for the development of more reliable vision–language models.
[1] T. Guan et al., “HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models,” Mar. 25, 2024, arXiv: arXiv:2310.14566. doi: 10.48550/arXiv.2310.14566.
[2] Y. Li, Y. Du, K. Zhou, J. Wang, W. X. Zhao, and J.-R. Wen, “Evaluating Object Hallucination in Large Vision-Language Models,” Oct. 26, 2023, arXiv: arXiv:2305.10355. doi: 10.48550/arXiv.2305.10355.
[3] A. Seth, D. Manocha, and C. Agarwal, “Towards a Systematic Evaluation of Hallucinations in Large-Vision Language Models,” Mar. 13, 2025, arXiv: arXiv:2412.20622. doi: 10.48550/arXiv.2412.20622.
[4] Z. Liu et al., “Safety Alignment for Vision Language Models,” May 22, 2024, arXiv: arXiv:2405.13581. doi: 10.48550/arXiv.2405.13581.
[5] W. Luo, S. Ma, X. Liu, X. Guo, and C. Xiao, “JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks,” Nov. 24, 2024, arXiv: arXiv:2404.03027. doi: 10.48550/arXiv.2404.03027.
[6] Z. Yang et al., “Distraction is All You Need for Multimodal Large Language Model Jailbreaking,” Jun. 17, 2025, arXiv: arXiv:2502.10794. doi: 10.48550/arXiv.2502.10794.
[7] C. Jiang et al., “Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models,” in Proceedings of the 32nd ACM International Conference on Multimedia, Melbourne VIC Australia: ACM, Oct. 2024, pp. 525–534. doi: 10.1145/3664647.3680576.
[8] Q. Cao, J. Cheng, X. Liang, and L. Lin, “VisDiaHalBench: A Visual Dialogue Benchmark For Diagnosing Hallucination in Large Vision-Language Models,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand: Association for Computational Linguistics, 2024, pp. 12161–12176. doi: 10.18653/v1/2024.acl-long.658.
[9] S. Schrodi, D. T. Hoffmann, M. Argus, V. Fischer, and T. Brox, “Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models,” Apr. 16, 2025, arXiv: arXiv:2404.07983. doi: 10.48550/arXiv.2404.07983.
[10] S. Yin et al., “Woodpecker: Hallucination Correction for Multimodal Large Language Models,” Sci. China Inf. Sci., vol. 67, no. 12, p. 220105, Dec. 2024, doi: 10.1007/s11432-024-4251-x.
[11] V. Rawte, A. Mishra, A. Sheth, and A. Das, “Defining and quantifying visual hallucinations in vision-language models,” in Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025), 2025, pp. 501–510.
[12] W. Liang, Y. Zhang, Y. Kwon, S. Yeung, and J. Zou, “Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning,” Oct. 19, 2022, arXiv: arXiv:2203.02053. doi: 10.48550/arXiv.2203.02053.
[13] C. Yi, Y.-H. He, D.-C. Zhan, and H.-J. Ye, “Bridge the Modality and Capability Gaps in Vision-Language Model Selection,” May 18, 2025, arXiv: arXiv:2403.13797. doi: 10.48550/arXiv.2403.13797.
[14] X. Liu, Y. Zhu, J. Gu, Y. Lan, C. Yang, and Y. Qiao, “معیار ایمنی چندوجهی: A Benchmark for Safety Evaluation of Multimodal Large Language Models,” Jun. 19, 2024, arXiv: arXiv:2311.17600. doi: 10.48550/arXiv.2311.17600.
[15] Z. Liu et al., “PSA-VLM: Enhancing Vision-Language Model Safety through Progressive Concept-Bottleneck-Driven Alignment,” 2024, arXiv. doi: 10.48550/ARXIV.2411.11543.
[16] P. Kaul et al., “THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models,” Apr. 03, 2025, arXiv: arXiv:2405.05256. doi: 10.48550/arXiv.2405.05256.
[17] S. Wei, X. Li, Y. Yao, and S. Yang, “A Novel Short-Memory Sequence-Based Model for Variable-Length Reading Recognition of Multi-Type Digital Instruments in Industrial Scenarios,” Algorithms, vol. 16, no. 4, p. 192, Mar. 2023, doi: 10.3390/a16040192.
[18] X. Fan, X. Wang, H. Wei, X. Zhang, and D. Zhao, “Toward a Stable, Fair, and Comprehensive Evaluation of Object Hallucination in Large Vision-Language Models,” in Advances in Neural Information Processing Systems 37, Vancouver, BC, Canada: Neural Information Processing Systems Foundation, Inc. (NeurIPS), 2024, pp. 111406–111431. doi: 10.52202/079017-3538.
[19] M. Ye-Bin, N. Hyeon-Woo, W. Choi, and T.-H. Oh, “BEAF: Observing BEfore-AFter Changes to Evaluate Hallucination in Vision-language Models,” Jul. 18, 2024, arXiv: arXiv:2407.13442. doi: 10.48550/arXiv.2407.13442.
[20] Y. Shu et al., “When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding,” Oct. 07, 2025, arXiv: arXiv:2506.05551. doi: 10.48550/arXiv.2506.05551.
[21] X. Liu, M. Luo, A. Chatterjee, H. Wei, C. Baral, and Y. Yang, “Investigating VLM Hallucination from a Cognitive Psychology Perspective: A First Step Toward Interpretation with Intriguing Observations,” ArXiv Prepr. ArXiv250703123, 2025.
[22] R. Gong, “Objective or biased? CEO overconfidence and journalists’ coverage,” Account. Finance, vol. 64, no. 4, pp. 4333–4357, Dec. 2024, doi: 10.1111/acfi.13311.
[23] W. Hu, S. Gu, Y. Wang, and R. Hong, “VideoJail: Exploiting Video-Modality Vulnerabilities for Jailbreak Attacks on Multimodal Large Language Models,” presented at the ICLR 2025 Workshop on Building Trust in Language Models and Applications, Mar. 2025. Accessed: Feb. 03, 2026. [Online]. Available: https://openreview.net/forum?id=fSAIDcPduZ
[24] H. Zhong et al., “Towards Visualization-of-Thought Jailbreak Attack against Large Visual Language Models,” presented at the The Thirty-ninth Annual Conference on Neural Information Processing Systems, Oct. 2025. Accessed: Feb. 03, 2026. [Online]. Available: https://openreview.net/forum?id=5P5YgohyBZ
[25] S. S. Ghosal et al., “Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment,” Jun. 14, 2025, arXiv: arXiv:2411.18688. doi: 10.48550/arXiv.2411.18688.
[26] L. He, Z. Chen, Z. Shi, T. Yu, J. Shao, and L. Sheng, “Systematic Reward Gap Optimization for Mitigating VLM Hallucinations,” Nov. 24, 2025, arXiv: arXiv:2411.17265. doi: 10.48550/arXiv.2411.17265.
[27] X. Lyu, B. Chen, L. Gao, J. Song, and H. T. Shen, “Alleviating Hallucinations in Large Vision-Language Models through Hallucination-Induced Optimization,” Apr. 01, 2025, arXiv: arXiv:2405.15356. doi: 10.48550/arXiv.2405.15356.
[28] Y. Shu, “Behavioural Finance and ESG: Current Landscape Summary,” in Proceedings of the 2025 3rd International Academic Conference on Management Innovation and Economic Development (MIED 2025), vol. 348, B. Siuta-Tokarska, A. Grigorescu, Md. M. Habib, and Y. Zhu, Eds., Dordrecht: Atlantis Press International BV, 2025, pp. 473–481. doi: 10.2991/978-94-6463-835-6_50.
[29] Z. Sun et al., “Aligning Large Multimodal Models with Factually Augmented RLHF,” Sep. 25, 2023, arXiv: arXiv:2309.14525. doi: 10.48550/arXiv.2309.14525.
[30] Y. Xie, G. Li, X. Xu, and M.-Y. Kan, “V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization,” Nov. 05, 2024, arXiv: arXiv:2411.02712. doi: 10.48550/arXiv.2411.02712.
[31] Z. Yang, X. Luo, D. Han, Y. Xu, and D. Li, “Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key,” Mar. 03, 2025, arXiv: arXiv:2501.09695. doi: 10.48550/arXiv.2501.09695.
[32] B. Chen, X. Lyu, L. Gao, J. Song, and H. T. Shen, “SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore Mechanism,” Dec. 03, 2025, arXiv: arXiv:2507.01513. doi: 10.48550/arXiv.2507.01513.
[33] E. Shayegani, Y. Dong, and N. Abu-Ghazaleh, “Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models,” Oct. 10, 2023, arXiv: arXiv:2307.14539. doi: 10.48550/arXiv.2307.14539.
[34] N. Jiang, A. Kachinthaya, S. Petryk, and Y. Gandelsman, “Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations,” Feb. 10, 2025, arXiv: arXiv:2410.02762. doi: 10.48550/arXiv.2410.02762.
[35] X. Chen et al., “Multi-Object Hallucination in Vision-Language Models,” Oct. 31, 2024, arXiv: arXiv:2407.06192. doi: 10.48550/arXiv.2407.06192.