Journal of Innovations in Social and Applied Sciences

A peer-reviewed, open-access journal for global research and innovation
Volume 1 (2025), Issue 2

Improving Visual Grounding and Clinical Reasoning in Medical Vision-Language Models

Uzair Iqbal Department of Software Engineering, University of Malaya
Hamid Wadood University of Alabama
Open Access — CC-BY-4.0
Download PDF Read Online

Keywords: Medical Vision-Language Models; Visual Grounding; Clinical Reasoning; Medical Visual Question Answering; Parameter-Efficient Fine-Tuning
Abstract

Medical vision-language models (VLMs) have increasingly been able to answer medical questions from images and perform multimodal reasoning; however, high accuracy on medical images does not necessarily mean predictions are made in clinically relevant regions. In this research, we examine the relationship between visual grounding and clinical reasoning by proposing an evidence-guided VLM framework in which visual features are explicitly localized and incorporated into the generation of answers to the questions. The proposed method implements a mostly frozen vision-language backbone network, a visual evidence module (text-only), and parameter-efficient adaptation using the SLAKE medical VQA dataset to reduce computational cost while maintaining multimodal representation ability. A composite objective is used to optimize the model, considering both the correctness of the answers and the consistency between the predicted evidence and the available semantic image annotations, as well as between the localized information and the clinical predictions. The framework is assessed with respect to traditional open-ended and closed-ended accuracy VQAs, spatial grounding accuracy VQAs, and evidence-consistency VQAs involving selective preservation or deletion of clinically relevant regions to quantify their effects on model predictions. This formulation allows not only for the assessment of the correctness of the medical VLM, but also the correctness of the visual evidence. The resulting framework will serve as a basis for computationally efficient medical VLMs with clinical predictions more closely grounded in the image evidence, and reasoning behavior assessed beyond the level of answer accuracy.

How to Cite
Uzair Iqbal, Hamid Wadood (2025). Improving Visual Grounding and Clinical Reasoning in Medical Vision-Language Models. Journal of Innovations in Social and Applied Sciences (JISAS), 1(2), 11-24.