The strongest version of the Google Gemini healthcare evidence story is simple enough to fit on a procurement slide: in a June 2026 Nature Medicine study, Gemini 3.1 Pro scored 97.4% on MedQA, ahead of OpenEvidence at 89.6% and UpToDate Expert AI at 88.4%; GPT-5.2 and Claude Opus 4.6 also outperformed the specialized clinical AI tools in the reported ranking.[1]
That is not a rounding error. It is a meaningful gap on a medical reasoning benchmark, and it deserves attention from health systems that have been paying for domain-specific tools on the assumption that clinical specialization buys safer or better answers. But the study does not turn a leaderboard into a purchasing order. It changes what should be tested locally; it does not decide what should be deployed.

| What the June 2026 evidence supports | What it does not settle |
|---|---|
| Frontier models performed very strongly on MedQA, with Gemini 3.1 Pro leading the reported comparison. | Whether public benchmark exposure inflated any model's score. |
| The reported pattern extended beyond MedQA to HealthBench and a NYU Langone clinical-query review. | Whether a model that answers more often is safer than one that refuses underspecified prompts. |
| No model in the study produced more harmful content or hallucinations than another. | Whether benchmark performance predicts effectiveness inside a live clinical workflow. |
| Specialized vendors can no longer assume their clinical branding is enough to win reasoning comparisons. | Whether any of these general-purpose models should be procured as clinical decision-support devices. |
The Result Is Too Strong To Dismiss
The study's main claim is not just that one frontier model beat one clinical tool on one exam-style test. The reported MedQA ranking put Gemini 3.1 Pro at 97.4%, above OpenEvidence at 89.6% and UpToDate Expert AI at 88.4%, while GPT-5.2 and Claude Opus 4.6 also ranked ahead of the specialized comparators.[1] A procurement team can reasonably treat that as evidence that the benchmark gap between general-purpose and clinical-branded systems has narrowed, and perhaps reversed, at least on the tasks measured.
The more interesting part is that the pattern reportedly held outside MedQA. The study also evaluated HealthBench and 100 real clinical queries from NYU Langone, using blinded review by 12 clinicians and 1,800 annotations.[1] That does not make the evaluation equivalent to bedside deployment, but it does move the evidence closer to the kinds of questions clinicians actually ask than a public multiple-choice bank alone.
The safety signal, as reported, also matters. The study found that no model produced more harmful content or hallucinations than another.[1] If that finding holds under scrutiny, it weakens a familiar defense of specialized tools: that they may score lower but behave materially more safely. In this study, that safety separation did not appear.
This is why the result should not be waved away as another flashy benchmark. It is credible enough to force a serious reassessment of assumptions about clinical AI vendors. A health system that has treated OpenEvidence or UpToDate Expert AI as categorically different from frontier-model products now needs better evidence than brand category, interface polish, or the phrase "built for clinicians."
It is also exactly the kind of result that can become dangerous when compressed. The version that belongs in governance review is not "Gemini beat clinical AI." It is: a June 2026 study found that frontier models, including Gemini 3.1 Pro, outperformed specialized clinical AI tools on MedQA, with similar reported ordering on HealthBench and a blinded review of 100 real clinical queries, while leaving unresolved questions about contamination, refusal behavior, and clinical actionability.[1]
Benchmark Exposure Comes First
OpenEvidence has formally requested that Nature Medicine retract the study, arguing that public benchmark questions may have appeared in frontier-model training data.[1] That objection cannot be treated as mere vendor wounded pride, even if the commercial stakes are obvious. If a model has seen benchmark items or close variants during training, its score may partly measure memorization or benchmark familiarity rather than clinical reasoning under uncertainty.
The uncomfortable detail is that the study authors themselves acknowledged the possibility. They wrote that frontier models "may have been exposed to MedQA or HealthBench during training."[1] That sentence does not invalidate the whole paper. It does, however, limit the purchasing inference a hospital can safely draw from the benchmark portion of the work.
The NYU Langone query review helps because real clinical queries are less likely to be replicated in public training corpora than widely used benchmark questions. But it does not erase the contamination issue for MedQA or HealthBench, and it does not fully answer whether the frontier models are better because they reasoned better, trained broader, saw similar tasks before, or benefited from all three.
For procurement, that distinction is not academic. A contaminated benchmark can still reveal that a model is broadly capable, but it cannot carry the same weight as a sealed local evaluation built from recent, institution-specific cases. If a health system is considering a vendor switch, the next test should use queries that were not public, were not vendor-selected, and were reviewed by clinicians who do not know which system produced which answer.

Refusal May Be A Feature
Wolters Kluwer's counterargument is different. The company did not merely say that UpToDate Expert AI should have scored higher. It argued that the tool's 19% refusal rate reflected safety awareness: declining to answer underspecified questions rather than producing an answer when the prompt lacked enough clinical context.[1]
That matters because many benchmark formats reward completion. If a question has a correct answer in the test key, a refusal looks like failure. In clinical work, refusal can be failure, but it can also be the right behavior. A tool that asks for age, pregnancy status, renal function, medication list, or acuity before recommending a next step may look less helpful in a static evaluation and more responsible in a real encounter.
The point is not that every refusal is wise. A system that refuses routine, well-specified questions wastes clinician time and may drive users to less governed tools. A system that answers everything may create a different problem: confident output in situations where the clinically safe move is to narrow the question first. The study's refusal-rate dispute should therefore push governance teams to classify refusals, not simply count them.
| Refusal type | Procurement interpretation |
|---|---|
| Refuses because the prompt is missing clinically necessary context. | Potentially appropriate guardrail; test whether the tool asks for the right missing information. |
| Refuses despite adequate clinical detail and a bounded request. | Workflow burden; may reduce usability and adoption. |
| Answers a materially underspecified prompt without caveats. | Potential safety concern even if the benchmark marks the answer correct. |
| Answers but clearly states assumptions and escalation conditions. | May be acceptable depending on local policy, specialty, and intended use. |
This is where a generic leaderboard becomes too blunt for clinical AI procurement. A CMIO does not only need to know which model gives the most correct answer to a known question. They need to know when the system recognizes that the question is not yet answerable. The Nature Medicine result raises that issue; it does not settle it.
Benchmarks Still Are Not Clinics
The broader evidence base gives procurement teams a reason to slow down even when a new benchmark result looks impressive. A 2025 npj Digital Medicine meta-analysis of 83 studies found that generative AI systems had an overall diagnostic accuracy of 52.1%.[2] The same analysis found no significant difference versus non-experts, with p=0.93, and significant inferiority versus expert physicians, with p=0.007.[2]
The meta-analysis also found that 76% of included studies had high risk of bias according to PROBAST.[2] That is the floor under the caution here. Much of the published evidence on generative AI diagnosis has been methodologically fragile, and performance in controlled question-answering settings has not reliably translated into a clear demonstration of clinical safety or effectiveness.
This should not be misused as a rebuttal to Gemini 3.1 Pro specifically. The 2025 meta-analysis largely reflects earlier generative AI evidence, including earlier Gemini-generation studies, not the exact June 2026 model comparison. It does not prove that Gemini 3.1 Pro is unsafe or ineffective. It does show why a high score on MedQA should not be treated as if it were a clinical outcomes trial.
MedQA and HealthBench can test medical knowledge, reasoning patterns, and response quality under curated conditions. A live deployment tests other things: whether the tool integrates into the EHR without creating copy-paste risk; whether it handles incomplete notes, conflicting labs, and time pressure; whether clinicians over-trust fluent explanations; whether escalation rules are followed; whether performance holds across specialties and patient populations; and whether the organization can audit what happened after a recommendation was shown.
Those are not decorative implementation questions. They determine who bears the consequence when the answer is wrong, incomplete, or used outside its intended scope.
What The Study Should Change In Procurement
The Nature Medicine study should make specialized clinical AI vendors work harder. If their products cost more, impose narrower workflows, or claim superior clinical fit, they should be able to show where that value appears: fewer unsafe completions, better refusal behavior, better source grounding, better specialty performance, better clinician acceptance, cleaner audit trails, or stronger local outcomes. A general-purpose frontier model scoring higher on medical benchmarks is a direct challenge to vague claims of clinical superiority.
But the same study should make buyers more demanding of frontier-model vendors, not less. A hospital should not move from "specialized tools are safer" to "frontier models are better" on the strength of a single publication cycle. The right conclusion is that both categories now need to be evaluated against the same local standard.
- Use contamination-aware test sets built from recent local clinical questions, not public benchmark items.
- Run blinded clinician review with explicit scoring for correctness, harmfulness, hallucination, source quality, and actionability.
- Separate refusal analysis from accuracy analysis, because a refusal can be either a defect or a safety behavior.
- Test inside the intended workflow, including handoffs, documentation burden, escalation rules, and auditability.
- Keep regulatory status separate from benchmark status; a strong medical reasoning score is not the same thing as clearance for a clinical decision-support use.
- Appraise evidence by specialty and use case, rather than assuming one model ranking generalizes across the health system.
This is also where internal evidence strategy matters. A system already building an AI review process can place the Nature Medicine result into a staged evidence pipeline rather than treating it as a standalone event; the broader question is how benchmark signals, local validation, and real-world monitoring fit together over time. That is the useful frame for Google's healthcare AI evidence pipeline, and it is the frame procurement teams need here.
A Procurement Verdict, Not A Model Forecast
The June 2026 Nature Medicine study is important because it makes a narrow but consequential claim credible: frontier models such as Gemini 3.1 Pro, GPT-5.2, and Claude Opus 4.6 may now outperform specialized clinical AI tools on several medical reasoning evaluations, including MedQA, HealthBench, and a blinded review of real NYU Langone clinical queries.[1]
That should trigger reassessment, not replacement. The unresolved contamination dispute limits confidence in public benchmark rankings. The UpToDate Expert AI refusal-rate argument shows that some behavior counted against a model may be safety-relevant rather than purely deficient. The wider diagnostic evidence base shows why benchmark performance has to be disciplined by clinical validation rather than promoted into deployment logic.[2]
The safe decision this evidence supports is local testing: contamination-aware evaluation, blinded clinician review, refusal-behavior analysis, workflow simulation, specialty-specific appraisal, and a clear separation between benchmark performance and regulated clinical use. The unsafe decision is the shortcut: MedQA rank becomes vendor switch, and the governance committee is asked to defend the missing steps later.
References
- ChatGPT, Gemini, Claude beat clinical AI tools: Study. Becker's Hospital Review.
- Diagnostic accuracy of generative artificial intelligence in medicine: a systematic review and meta-analysis. npj Digital Medicine. 2025.