跳至主要内容
临床试验/NCT07281066
NCT07281066已完成不适用

Evaluating ChatGPT-4o, Gemini and Claude 3.7 in Endodontic Diagnostics: A Prospective Clinical Study

Marmara University1 个研究点 分布在 1 个国家目标入组 120 人开始时间: 2025年7月7日最近更新:

试验速览

阶段
不适用
状态
已完成
入组人数
120
试验地点
1
主要终点
Clinician Diagnosis Accuracy Based on Paper-Based History and Periapical Radiograph

研究概览

简要总结

The goal of this prospective observational study is to evaluate the ability of three large language models (ChatGPT-4o, Gemini Advanced, and Claude 3.7) to support diagnosis and treatment decision-making in adult patients presenting with common endodontic conditions.

The main questions the study aims to answer are:

Can LLMs accurately determine the endodontic diagnosis when provided with structured clinical information and periapical radiographs?

Can LLMs propose appropriate treatment plans comparable to decisions made by endodontic specialists?

To answer these questions, researchers will compare the diagnostic and treatment accuracy of three AI models using a consensus diagnosis from endodontic specialists as the reference standard.

Participants will:

Receive routine endodontic examination and periapical radiographs as part of standard clinical care.

Have their anonymized clinical histories and radiographs entered into the three AI models.

Not interact directly with any AI system; all evaluations will be performed by the research team.

This study aims to understand how large language models perform under real-world clinical conditions and whether these systems may play a supportive role in endodontic diagnostics in the future.

详细描述

This prospective observational study aims to evaluate the real-time diagnostic and treatment decision-making performance of three large language models-ChatGPT-4o, Gemini Advanced, and Claude 3.7-in an endodontic clinical setting. A total of 120 patients presenting to the endodontic clinic were examined, and detailed medical/dental histories, clinical findings, and periapical radiographs were collected. Each anonymized case was then presented to the three LLMs using a standardized prompt asking for the diagnosis and the appropriate treatment plan.

All models were used in their default multimodal configurations without enabling web-search functions, plug-ins, or external data retrieval. Each question was submitted only once in isolated chat sessions to prevent memory carry-over. Responses were saved verbatim and compared with the reference diagnoses and treatment plans established by a panel of endodontic specialists.

This study was designed to mimic real-world clinical conditions as closely as possible, providing a realistic assessment of how these systems might perform when used by clinicians in everyday practice. Understanding their capabilities and limitations in authentic clinical scenarios is essential, as LLMs are expected to play an increasingly vital role in future dental care particularly in decision support, triage, and patient education. By identifying where these models perform well and where they fall short, this research aims to inform safe and effective clinical integration as LLM technologies continue to advance.

研究设计

研究类型
Observational
观察模型
Case Only
时间视角
Prospective

入排标准

年龄范围
18 Years 至 65 Years(Adult, Older Adult)
性别
All
接受健康志愿者

入选标准

  • Adult patients (≥18 years old) presenting to or referred to the Endodontic Clinic.
  • Patients with a clinically verified endodontic condition requiring diagnosis and treatment planning.
  • Patients who agreed to participate and provided informed consent.
  • Patients for whom a complete paper-based medical/dental history and periapical radiograph were obtained during the clinical visit.

排除标准

  • Exclusion Criteria
  • Patients who declined participation or did not provide informed consent.
  • Pediatric patients (<18 years old) referred to the Pediatric Dentistry Clinic.
  • Patients attending the clinic with non-endodontic complaints (e.g., post-extraction alveolitis, third-molar extraction problems).
  • Cases with incomplete clinical information or missing radiographs.
  • Patients unable to undergo standard endodontic examination procedures.

结局指标

主要结局

Clinician Diagnosis Accuracy Based on Paper-Based History and Periapical Radiograph

时间窗: 7 july-5 august

Assessment of the diagnostic decision made by endodontic clinicians after reviewing a paper-based patient history form and a standardized periapical radiograph. Accuracy is determined by comparing the clinician's diagnosis with the consensus diagnosis established by three independent endodontic specialists. Data will be collected for all 120 patients at the time of initial clinical evaluation.

次要结局

  • LLM-Generated Diagnosis and Treatment Planning Performance(august-september)

研究者

申办方类型
Other
责任方
Sponsor

研究点 (1)

Loading locations...

相似试验

LLM Performance in Endodontic Diagnostics | 临床试验