Study Finds General-Purpose AI Outperforms Clinical Tools

A doctor in a lab coat holds a glowing orb that says AI
Photo credit: Dreamstime Photos

A growing number of specialized clinical artificial intelligence tools are being introduced into medical practice, often with claims that they perform better than general-purpose large language models.

 

A recent study published in Nature evaluated whether that advantage holds when these systems are tested directly. Across three rigorous evaluations, the study found that the frontier models outperformed the specialized clinical tools.

Methodology

The study used a three-part evaluation framework.

 

First, the authors randomly sampled 500 U.S. Medical Licensing Examination-style questions from MedQA. These questions were used to test medical knowledge.

 

Second, they sampled 500 single-turn prompts from HealthBench. This benchmark was used to assess how well model responses aligned with clinician expectations across five axes: accuracy, completeness, communication quality, context awareness and instruction following. Responses were also grouped into seven themes, including emergency referrals, context seeking, global health and responding under uncertainty. HealthBench responses were graded by a panel of three large language model judges: Claude Opus 4.6, Gemini 3.1 Pro Preview and GPT-5.2.

 

Third, the authors created a real clinical queries benchmark from 100 de-identified physician queries submitted to a Health Insurance Portability and Accountability Act-compliant GPT instance at NYU Langone. Each query was submitted to six systems: GPT-5.2, Gemini 3.1 Pro, Claude Opus 4.6, OpenEvidence, UpToDate Expert AI and Google Search AI Overview.

 

Twelve U.S. clinicians, blinded to model identity, reviewed the responses. They rated each response on four dimensions: clinical correctness, completeness, safety or harm avoidance and clarity, using a one-to-four scale.

 

Results

MedQA

General-purpose frontier models scored higher than the clinical AI tools on the 500 MedQA questions.

 

Gemini had the highest accuracy at 97.4%, followed by GPT at 94.2% and Claude at 90.2%. OpenEvidence scored 89.6% and UpToDate Expert AI scored 88.4%.

 

Gemini outperformed all other models. GPT also outperformed OpenEvidence, UpToDate and Claude.

HealthBench

On HealthBench, GPT had the highest score at 88.0 out of 100. Gemini scored 79.3 and Claude scored 77.0. The clinical tools scored lower: OpenEvidence at 62.6 and UpToDate at 61.3.

 

GPT outperformed all other models. The two clinical tools did not differ significantly from one another. In the theme-level analysis, GPT ranked first or tied for first in all seven categories. OpenEvidence and UpToDate ranked lowest or tied for lowest in all seven categories.

Real clinical queries

In the blinded review of 100 real clinical queries, the models separated into two performance tiers.

 

The first tier consisted of the frontier models:

  • Gemini: 3.62 mean aggregate rating
  • GPT: 3.54
  • Claude: 3.52

The second tier consisted of the clinical tools and Google Search AI Overview:

  • Google AI Overview: 3.27
  • OpenEvidence: 3.24
  • UpToDate Expert AI: 3.17

 

There were no significant differences within each tier, but all significant pairwise differences were between the tiers. According to the authors, this means frontier models outperformed the clinical tools on most individual questions, not only in average scores.

 

After adjustment for rater leniency, the clinical AI tools, including Google AI Overview, had 49%-87% lower odds of receiving a higher rating than Gemini. In a sensitivity model, this translated to scores that were 0.36 to 0.44 points lower on the one-to-four scale.

 

The same two-tier pattern held across all four evaluation dimensions. The models differed most on clarity and least on clinical correctness. OpenEvidence had the lowest clarity score, which the authors interpret as a communication weakness rather than a knowledge weakness.

 

The authors also note qualitative patterns in the responses. Incomplete clinical content, safety-critical omissions and disorganized answers were common, particularly for OpenEvidence and Google AI Overview.

 

UpToDate Expert AI had the highest refusal rate at 19%. That was higher than the refusal rates for the other systems, which ranged from 1%-6%.

 

On safety outcomes, the study did not find statistically significant differences among the models in harmful content or hallucinations. All 12 clinician raters showed similar ranking patterns, consistently placing the frontier models above the clinical tools.

What this means for eyecare providers

For eyecare providers, the main takeaway is not that one category of tool should be adopted broadly. Performance claims for specialized clinical AI systems should be independently tested rather than assumed.

 

This study found that the specialized clinical tools did not outperform the frontier models in medical knowledge, clinician-aligned evaluation or blinded review of real physician queries. In the real clinical query benchmark, the clinical tools performed similarly to Google Search AI Overview. For providers considering AI tools in ophthalmology or optometry workflows, that suggests brand positioning as a clinical product does not, by itself, establish superior performance.

 

The article also points to practical caution. The authors did not assess response latency or citation quality, both of which could affect clinical usability. They also note that performance in this area is evolving quickly, and they do not claim the ranking will remain fixed over time. They further state that deeply subspecialized tasks may still benefit from more sophisticated domain-specific adaptation.

 

For eyecare settings, the study supports careful local evaluation of AI tools on real use cases before integration into practice. That could include testing how a system handles specialty-specific clinical questions, communication clarity, completeness and refusal behavior in workflows relevant to eye care. The article’s broader point is that independent, real-world evaluation matters more than product category alone.

 

Learn more about this study on our Real Talk podcast here

 

Read more of the latest AI in Eye Care news here

Author

Leave a Reply

Your email address will not be published. Required fields are marked *