Nature Medicine Benchmark Study Finds General-Purpose LLMs Outperform FDA-Cleared Clinical AI Tools Across All Medical Question Types
核心洞察
A June 2026 Nature Medicine study found OpenAI's GPT-5.2, Google's Gemini 3.1 Pro, and Anthropic's Claude Opus 4.6 outperformed FDA-cleared clinical AI tools OpenEvidence (搜索) and UpToDate Expert AI (搜索) on every medical benchmark tested.
On real clinical queries from practicing physicians, frontier LLMs formed a clear top performance tier (mean scores 3.52–3.62) while clinical tools and Google AI Overview lagged behind (3.17–3.27), with 49–87% lower odds of receiving higher ratings.
The study exposes a regulatory gap: FDA clearance assesses whether a device meets sponsor-defined specifications, not whether it performs better than freely available unregulated alternatives that physicians are already using.
On June 23, 2026, Nature Medicine published a landmark head-to-head benchmarking study that delivers an unambiguous verdict: general-purpose large language models from OpenAI, Google, and Anthropic systematically outperform FDA-cleared clinical AI tools—OpenEvidence (搜索) and Wolters Kluwer (搜索)'s UpToDate Expert AI (搜索)—across every medical benchmark tested, including real-world physician queries submitted at the point of care.
The study, conducted by researchers at NYU Langone Health, evaluated six systems across three distinct stages: 500 US Medical Licensing Examination-style MedQA questions, 500 HealthBench items assessing alignment with expert clinicians, and 100 real clinical queries (RCQ) drawn from physician use during live clinical deployment. The RCQ stage underwent randomized, blinded review by 12 US clinicians, producing 1,800 model–question annotations.
Frontier Models Dominate Across All Benchmarks
On MedQA questions, Google's Gemini 3.1 Pro Preview achieved the highest accuracy at 97.4% (95% CI 95.6%–98.5%), followed by OpenAI's GPT-5.2 at 94.2% (91.8%–95.9%) and Anthropic's Claude Opus 4.6 at 90.2% (87.3%–92.5%). The clinical tools scored notably lower: OpenEvidence (搜索) achieved 89.6% (86.6%–92.0%) and UpToDate Expert AI (搜索) reached 88.4% (85.3%–90.9%). Gemini outperformed all other models with statistical significance (McNemar P < 1 × 10⁻⁴ versus OpenEvidence, UpToDate, and Claude; P = 0.02 versus GPT).
On HealthBench, GPT-5.2 scored highest at 88.0 (95% CI 85.9–90.1), followed by Gemini at 79.3 (76.6–81.9) and Claude at 77.0 (74.2–79.9). Both clinical tools scored substantially lower—OpenEvidence (搜索) at 62.6 (59.3–65.9) and UpToDate at 61.3 (58.0–64.6)—with GPT outperforming all other models (Wilcoxon P < 10⁻⁹). In theme-level analysis, GPT ranked first or tied for first in all seven categories, while OpenEvidence and UpToDate ranked lowest or tied for lowest in all seven.
Real-World Clinical Queries Reveal a Two-Tier Performance Structure
The RCQ benchmark, which the authors designate as the study's primary evidence, sampled 100 anonymous clinician queries from NYU Langone's HIPAA-compliant GPT instance. Twelve blinded clinicians scored responses across four dimensions: clinical correctness, completeness, safety/harm avoidance, and clarity.
The six models differed significantly (Friedman P < 10⁻⁹), with two clear performance tiers emerging. Frontier LLMs formed the top tier: Gemini (mean aggregate 3.62; 95% CI 3.56–3.68), GPT (3.54; 3.47–3.61), and Claude (3.52; 3.44–3.59), with no significant differences between them. Clinical tools and Google AI Overview followed in the second tier: OpenEvidence (搜索) (3.24; 3.17–3.32), UpToDate AI (3.17; 3.09–3.25), and Google AI Overview (3.27; 3.18–3.35).
After adjusting for rater leniency, clinical AI tools had 49–87% lower odds of receiving a higher rating than Gemini (odds ratio 0.13–0.51; all P < 0.0001). In a sensitivity linear mixed model, this corresponded to 0.36–0.44 points lower on the 1–4-point scale (all P < 0.0001). Notably, Google AI Overview—an auto-enabled search feature routinely encountered by clinicians—matched or exceeded the performance of both FDA-cleared clinical tools across all dimensions.
Safety and Refusal Patterns
Safety outcomes did not differ across models: none produced more harmful content (Cochran's Q = 4.00, P = 0.55) or hallucinations (Q = 5.00, P = 0.42) than the others. However, UpToDate AI refused 19% of queries, significantly more than all other models (1–3%; P < 0.01) except Google AI Overview (6%; P = 0.10). OpenEvidence (搜索) scored lowest on clarity (mean 2.84), suggesting its weakness was communication rather than knowledge.
The Regulatory Gap Exposed
The study surfaces a structural issue in the current FDA regulatory framework. FDA's December 2024 final guidance on Predetermined Change Control Plans (PCCPs) for AI-Enabled Device Software Functions governs how a cleared tool can change post-clearance—not whether the cleared tool's baseline performance was adequate relative to unregulated alternatives performing the same clinical function.
The FDA's AI-Enabled Medical Devices List documents regulatory process completion but does not document comparative clinical performance against general-purpose models. When a cleared tool underperforms an uncleared one on the same physician query set, the clearance record offers no signal of that performance gap.
The FDA's February 2025 warning letter to Exer Labs, Inc. illustrates where current enforcement focus sits: on tools marketed without any clearance. This does not address the category the Nature Medicine study reveals—tools that have clearance but whose performance falls below what unregulated foundation models now deliver.
Implications for Clinical Trial Sponsors
For sponsors running AI-assisted clinical trials, the findings carry immediate operational significance. A CRO or eClinical platform integrating a cleared clinical AI tool for protocol deviation flagging, safety signal detection, or eCOA response interpretation is relying on clearance as a proxy for fitness-for-purpose. The benchmark data indicates that proxy may be unreliable.
The EMA's 2024 reflection paper on AI/ML in clinical development adopted a risk-tiered approach, implying that high-risk tools influencing trial primary endpoints or safety reporting should demonstrate superiority—or at minimum non-inferiority—to available alternatives. The Nature Medicine study introduces a documentation question not currently covered by any FDA guidance: if a cleared tool underperforms a general-purpose model on query types encountered in a trial, what is the sponsor's evidentiary obligation to document that performance differential in the IND or trial master file?
Study Limitations and Context
The authors acknowledge several limitations. Clinical tools lack public APIs, so they were queried through browser interfaces, which may have introduced differences in hidden prompts, retrieval behavior, and output formatting. Standardized benchmarks carry known risks of data leakage, though the RCQ benchmark is free from training-set contamination. HealthBench, as an OpenAI-developed benchmark, may be influenced by potential benchmark–developer overlap.
The study did not assess response latency or citation quality—factors important for real-world clinical deployment. The authors emphasize that results should be interpreted as "a snapshot of a rapidly evolving landscape rather than a permanent ordering of approaches," noting that deeply subspecialized medical tasks may favor more sophisticated, domain-specific adaptation.
The AMA's 2024 physician AI survey found that 66% of physicians currently use AI in practice, up from 38% in 2023, with 68% reporting perceived definite advantage. This adoption rate means physicians are already making treatment-adjacent decisions with AI assistance, making the regulatory question of comparative performance increasingly urgent.
FDA's Digital Health Center of Excellence has indicated it expects to issue updated guidance on clinical decision support software categorization in 2026. How that guidance defines "clinical decision support" relative to general-purpose LLMs performing the same function will determine whether the documented performance gap becomes a regulatory classification question or remains an unresolved sponsor risk.
