How Well Do Multimodal LLMs Reason in Eye Care?

A doctor holds a tablet with a ChatGPT logo above
Photo credit: Dreamstime Photos

Large language models (LLMS) have performed well on text-based medical exams, but ophthalmology often requires clinicians to interpret both written case details and images. That makes multimodal systems, which combine language and image analysis, especially relevant to the field.

 

In a recent study published in British Journal of Ophthalmology, researchers evaluated how three vision-language LLMs handled bilingual ophthalmology questions that paired clinical vignettes with images, with a focus on both answer accuracy and the quality of the models’ reasoning.

 

Methodology

Researchers assessed three multimodal large language models: CLM-V, ChatGPT-5 and MiniCPM-V 4.5. The models were tested on 316 bilingual ophthalmology questions, including 175 English single-choice questions from the Basic and Clinical Science Course and 141 Chinese multiple-choice questions from senior professional title exams.

 

The questions covered cornea, uvea, glaucoma, retina and orbit. Each question included a clinical vignette and an image.

 

The researchers tested each model under two conditions: with reasoning-enabled prompts and with reasoning-disabled prompts. They measured accuracy against reference standards. They also evaluated reasoning quality through automated rubric scoring that examined accuracy, data synthesis, logic, option analysis and safety. Human experts also reviewed the reasoning, and four cases were analyzed qualitatively.

 

Results

Reasoning-enabled prompting was associated with higher mean artificial intelligence-assisted total scores across all three models in both datasets.

 

In the English dataset, CLM-V’s mean score increased from 14.97 to 16.07, ChatGPT-5’s from 20.77 to 23.97, and MiniCPM-V 4.5’s from 10.83 to 12.60.

 

In the Chinese dataset, CLM-V’s score increased from 9.03 to 10.27, ChatGPT-5’s from 19.95 to 22.00, and MiniCPM-V 4.5’s from 11.05 to 13.30.

 

Human evaluation showed substantial inter-rater agreement, with a kappa value of 0.87. In those evaluations, ChatGPT-5 ranked highest.

 

The qualitative case analyses suggested that reasoning-enabled outputs were often more clinically interpretable. However, the study said the size of that benefit varied by model and by dataset.

 

What this means for eyecare providers

The findings suggest multimodal large language models may have value in ophthalmic question-answering, particularly when prompts are designed to elicit reasoning. For eyecare providers, that may be most relevant in educational settings or in situations where interpretability matters alongside answer selection.

 

At the same time, the study pointed to limitations in subspecialty robustness and image interpretation. That means performance may not be consistent across all ophthalmic domains or question types. The authors also emphasized the need for rigorous reasoning evaluation before these systems are used in educational or clinical applications.

 

 

Read more of the latest AI in Eye Care News here

Author

Leave a Reply

Your email address will not be published. Required fields are marked *