Autonomous Medical AI Agent MIRA Outperforms Physicians in Simulated EHR Diagnostic Cases
核心洞察
MIRA (搜索), an autonomous AI agent operating within a sandboxed EHR environment, achieved 87.8% diagnostic accuracy compared to 78.1% for board-certified physicians (p < 0.001) in a matched 311-case comparison.
The agent demonstrated perfect recall for critical admission decisions in pneumonia (搜索) and pulmonary embolism (搜索), with zero high-severity drug-drug interactions and zero renal dosing incompatibilities across 468 prescriptions.
MIRA (搜索) was evaluated on 574 real-world emergency department cases from the MIMIC-IV database spanning eight diagnoses, using 11 specialized digital tools with over 85,000 operational choices.
A new autonomous artificial intelligence agent designed to operate within electronic health record (EHR) systems has demonstrated the ability to outperform experienced physicians in simulated clinical scenarios, marking a significant step toward AI systems that can independently manage end-to-end diagnostic workflows.
The agent, called MIRA (搜索), was developed to translate clinical reasoning into structured EHR actions — ordering tests, synthesizing results, and producing diagnoses and treatment plans — within a sandboxed environment compliant with HL7 FHIR standards. In a study published in the journal Nature, MIRA achieved 88.9% diagnostic accuracy across 574 real-world emergency department cases drawn from the MIMIC-IV database, and 87.8% accuracy in a matched 311-case comparison against human physicians.
Performance against human clinicians
In the matched comparison, board-certified physicians reached an average accuracy of 78.1% (p < 0.001), while a mixed-seniority team consisting of four residents and two board-certified doctors averaged 71.1% (p < 0.001). The cases encompassed eight distinct diagnoses across surgery (appendicitis (搜索)), internal medicine (pneumonia (搜索)), and oncology (pancreatic cancer (搜索)).
MIRA (搜索) excelled particularly at identifying appendicitis (搜索) and pancreatitis (搜索), achieving a perfect 100% recall for laparoscopic appendectomies. For pancreatic cancer (搜索), its diagnostic performance was equivalent to that of board-certified physicians, while pneumonia (搜索) and urinary tract infections remained more challenging for the AI system.
Notably, MIRA (搜索) did not achieve its superior accuracy through indiscriminate test ordering. While the agent requested a broader, more comprehensive set of individual blood parameters than human doctors, its overall test selection remained well below historical dataset baselines. The AI model successfully avoided systematic over-ordering of high-cost radiological imaging, matching or exceeding physicians in overall resource-alignment metrics.
Safety evaluation results
An independent, blinded medical review of 56 patient-level outputs and a separate assessment of 468 prescriptions written by MIRA (搜索) found that the agent caused zero high-severity drug-drug interactions, zero renal dosing incompatibilities, and zero medication-allergy mismatches. Route specification was the weakest prescription field, at 97% correctness.
When making critical hospital admission decisions for pneumonia (搜索) and pulmonary embolism (搜索), MIRA (搜索) achieved a perfect recall score of 1.00, indicating that the AI tool never missed a single patient who required inpatient care. However, the pulmonary embolism analysis suggested a tendency toward over-admission, reflecting what researchers described as a cautious disposition strategy.
Architecture and methodology
MIRA (搜索) operates within an EHR sandbox using a suite of tools to simulate clinical workflows. It can order tests, synthesize results, and produce diagnoses and treatment plans while interacting through chat with a patient AI agent grounded in documented histories of present illness extracted from retrospective notes of real cases. The agent navigated its clinical environment using 11 specialized digital tools with more than 85,000 operational choices.
The patient AI agent was instructed to respond to questions posed by MIRA (搜索) or its human counterparts solely based on authentic clinical histories, while resisting adversarial attempts to trick it into prematurely leaking information. The authors noted, however, that simulated patient speech may be more structured than real emergency department conversations.
Limitations and future directions
The researchers caution that MIRA (搜索) and similar AI agents are not replacements for expert human staff. The model did not achieve 100% perfection in all treatment choices, such as specific antibiotic selections, highlighting the ongoing need for strict human supervision and patient-level safeguards.
Future model iterations may improve performance by incorporating evidence from retrieval-based support, stronger governance, and prospective real-world validation before any clinical deployment. The study underscores the safeguards needed before autonomous AI reaches real patient care settings.
