Head-to-Head Evaluation of ChatGPT 4o, GPT-5, and DeepSeek for Structured Extraction, Toric IOL Recommendation, and Refractive Prediction
试验速览
- 阶段
- 不适用
- 状态
- 招募中
- 发起方
- 入组人数
- 100
- 试验地点
- 1
- 主要终点
- Refractive prediction error for sphere
研究概览
简要总结
We conducted a single-center, retrospective observational study to evaluate large language models (ChatGPT 4o, GPT-5, DeepSeek) for automated interpretation of de-identified IOLMaster 700 reports provided as raster images. Models produced structured biometric extraction, toric IOL recommendation, and refractive predictions (sphere, cylinder, axis). Primary outcomes included parameter-level agreement and refractive error metrics; secondary outcomes included decision-support performance for toric IOL selection and agreement on ordered T-codes. No clinical intervention was performed.
详细描述
This study compares three large language models accessed in their native configurations, without fine-tuning or external tools. For each examination, the original IOLMaster 700 report image was supplied without manual annotation or pre-processing. A standardized instruction required: (i) structured extraction of AL, ACD, LT, WTW, K1/K2 and axes, ΔK, TK1/TK2 and axes, and ΔTK; (ii) binary toric candidacy and T-code according to institutional ALCON mapping; and (iii) refractive recommendations (sphere, cylinder, implantation axis). Each model generated three independent outputs per case. De-identification and IRB oversight (waiver of consent) were implemented according to institutional policy. The unit of enrollment is participants (n=54), with outcomes analyzed per eye (162 eyes) and per model generation where applicable.
研究设计
- 研究类型
- Observational
- 观察模型
- Case Only
- 时间视角
- Retrospective
入排标准
- 年龄范围
- 18 Years 至 —(Adult, Older Adult)
- 性别
- All
- 接受健康志愿者
- 否
入选标准
- •postoperative corrected distance visual acuity (CDVA) of 0.10 logMAR or better -an absolute IOL rotational stability of less than 10∘ at the 1-month follow-up examination
排除标准
- •incomplete biometric data on the examination report;
- •a history of previous ocular surgery or ocular trauma
- •the occurrence of intraoperative complications, such as an anterior capsular tear or posterior capsular rupture
- •the development of significant postoperative complications, including but not limited to severe intraocular infection or inadequate pupillary dilation.
结局指标
主要结局
Refractive prediction error for sphere
时间窗: At index examination
Mean absolute error (MAE, diopters) of model-predicted sphere versus clinical reference
Cohen's kappa with 95% CIs between model
时间窗: At index examination (single time point)
Cohen's kappa with 95% CIs between model outputs and clinician-validated reference for per-parameter
次要结局
- Cylinder prediction error(At index examination)
- Axis prediction error(At index examination)
研究者
Jin Yang
Chief Physician
Eye & ENT Hospital of Fudan University
