ChatGPT Outperforms Physicians in Diagnostic Study, Raising Questions About AI's Role in Medicine
核心洞察
A Harvard-Stanford study published in Science found that OpenAI (搜索)'s o1 model matched or exceeded experienced physicians on diagnostic tasks using real emergency-room cases.
Researchers caution that the study does not prove AI is ready for standard medical practice, while unvetted AI tools are already proliferating across U.S. healthcare.
Diagnostic errors kill or permanently disable nearly 800,000 Americans annually, and poorly controlled chronic diseases contribute to 50% of heart attacks, strokes, and kidney failures.
A landmark study published in the journal Science has demonstrated that OpenAI (搜索)'s o1 preview model can match or exceed the diagnostic performance of experienced physicians, marking what some observers describe as a potential inflection point for artificial intelligence in medicine. The research, led primarily by Harvard and Stanford investigators, pitted ChatGPT (搜索) against hundreds of physicians in a diagnostic obstacle course involving written medical mysteries and information from real-world patients. The AI won.
The study tested the model on difficult clinical tasks, including 76 real emergency-room cases at Beth Israel Deaconess Medical Center in Boston. At three stages of care — initial triage, first physician contact, and admission to the hospital — the model matched or exceeded the performance of experienced physicians on text-based diagnostic and clinical management tasks.
Yet the lead author struck a notably cautious tone. "I get a little bit queasy about how some of these results might be used," said Adam Rodman at a press conference ahead of the publication. He emphasized that the work amounted to an academic exercise and did not prove that ChatGPT (搜索) or any other AI tool was ready to become a standard part of medical practice.
AI Already Embedded in Clinical Workflows
Despite expert caution, generative AI has already permeated the U.S. health-care system. Clinicians at many institutions now receive regular emails announcing the availability of "AI-powered clinical reasoning tools," none of which has been approved for medical use by the FDA. This enthusiasm feels unprecedented in a field that still relies on pagers and fax machines — a reflection of medicine's traditionally safety-focused culture, where any ill-timed glitch has the potential to turn deadly.
Clinicians are now "allowed — encouraged, even — to run wild with the latest software, guided by a generic warning that 'AI can make mistakes,'" according to one pathologist's account. Those mistakes can be consequential. A randomized trial published in NEJM AI found that intentionally erroneous output from an AI model can easily lead doctors astray. A separate study led by researchers at Mount Sinai suggested that chatbots may fail to alert users to potential medical emergencies.
The Regulatory Gap
Part of the problem lies in how health-related AI products are classified. If a software package intended for physicians is categorized as a "clinical decision support tool" rather than a medical device, it typically avoids FDA oversight. To qualify, an AI-powered app generally must rely on existing medical literature, avoid analyzing medical scans or images, explain its reasoning, and leave diagnosis and treatment up to a physician. Most generative-AI products used by doctors today appear to meet these criteria.
Consumer-wellness apps enjoy similar latitude so long as they are intended for "maintaining or encouraging a healthy lifestyle" rather than diagnosing or treating specific conditions. Microsoft (搜索), OpenAI (搜索), Anthropic (搜索), and xAI all warn users that their health-related chatbots are not meant to provide medical care. In practice, however, the distinction is not always clear. Elon Musk encourages people to use his Grok chatbot to generate second medical opinions and interpretations of X-ray and MRI images. A marketing video for ChatGPT (搜索) Health shows the app reassuring people that their lab results are in a healthy range and encouraging them to continue taking cholesterol medication.
Hims & Hers (搜索) has launched a product called Labs AI, which helps interpret results from "up to 130 biomarker tests" and provides a "deep, personalized, and actionable analysis on whole body health, risks, and patterns." When contacted, Patrick Carroll, chief medical officer of Hims & Hers, stated that Labs AI does not diagnose or recommend treatment: "That responsibility belongs to clinicians, and Labs is designed to reinforce that boundary." Dominic King, vice president of health at Microsoft (搜索) AI, said its Copilot app provides "helpful information and support for conversations with clinicians" and not "a single, firm diagnosis."
The Human Cost of the Status Quo
The stakes of getting AI integration right — or wrong — are enormous. Diagnostic errors are estimated to kill or permanently disable nearly 800,000 Americans each year, according to research compiled by the Johns Hopkins Armstrong Institute Center for Diagnostic Excellence. Chronic diseases like hypertension (搜索) and diabetes (搜索) remain poorly controlled for millions of Americans, contributing to as many as 50% of all heart attacks, strokes, and kidney failures, per CDC data.
A survey from Notable found more than 60% of patients skipped a doctor visit in the past year because scheduling was too much of a hassle, and nearly half of U.S. adults worry they cannot afford needed care in 2026, according to a West Health-Gallup survey.
A New Paradigm for AI Evaluation
One idea circulating in the medical literature is to stop treating AI products as if they were merely standard medical devices. Given their humanlike ability to learn new information and tailor answers to individual patients, medical AIs may function more like doctors than defibrillators — so perhaps they should be evaluated in the same way physicians are. Instead of requiring FDA approval for each function, a chatbot might be asked to pass a medical-licensing exam and undergo a period of supervision akin to a medical residency.
Haider Warraich, a cardiologist and program manager at the Advanced Research Projects Agency for Health (ARPA-H), is leading a major effort to get medical chatbots approved through the traditional FDA pathway. His agency is funding the development of an AI tool tailored for heart conditions and intends to send it through a full FDA-authorization process, with the goal of enabling the chatbot to safely evaluate and treat patients without physician involvement. Rodman praised this approach but warned that the process will take years, during which time a plethora of new health AIs will slip into the market with little scrutiny.
The Uber Parallel
The emergence of today's AI health products has drawn comparisons to the rise of ride-sharing services in the 2010s. The taxi industry was heavily regulated, yet Uber and Lyft skirted those rules to acquire a critical mass of users rapidly. Governments eventually had little choice but to adjust their laws to match the new status quo. The same pattern could play out in medicine: regulations meant to ensure safety and effectiveness may either remain in force or be weakened to clear the path for tools that everyone is already using.
The trajectory of large language models underscores the urgency. In late 2022, Google's first medical AI model, Med-PaLM, achieved a passing score of 67% on the U.S. medical licensing exam. Four months later, Med-PaLM 2 scored at an expert doctor level of 87%. Assuming the next three years resemble the last three, GenAI tools will be eight to 16 times more powerful and clinically useful than they are today.
Rather than viewing generative AI mainly as a tool to decrease head count, proponents argue that payers should recognize the hundreds of billions of dollars that could be saved through better health and fewer life-threatening medical problems. The first step, they contend, is acknowledging the preventable deaths caused each year by today's medical system — and then embracing the ways that GenAI can support clinicians, empower patients, and save lives.
