Agreement Between Large Language Models and Faculty Assessment in the Evaluation of Clinical Reasoning Case Examinations in Undergraduate Physiotherapy Education: A Comparative Reliability Study
试验速览
- 阶段
- 不适用
- 状态
- 尚未招募
- 发起方
- 入组人数
- 65
- 试验地点
- 1
- 主要终点
- Agreement between LLM global scores and the faculty reference global score
研究概览
简要总结
This study evaluates whether large language models (LLMs) can reliably assess written clinical-reasoning case examinations completed by undergraduate physiotherapy students, compared with faculty assessment. In the course "Specific Methods in Physiotherapy" (third year of the Physiotherapy Degree), students solve complex clinical cases that require clinical reasoning, technical knowledge, and therapeutic decision-making. These cases are traditionally graded by faculty, a time-consuming process that may show inter-rater variability.
A set of de-identified student case examinations will be assessed using the rubric currently applied in the course, which covers clarity and structure of clinical reasoning, integration of the biopsychosocial model (ICF and APTA frameworks), accuracy in identifying pain mechanisms, coherence between diagnosis, hypotheses, and treatment, originality and depth of analysis, and professional writing. Each examination will be scored independently by three LLMs (for example, Claude, ChatGPT, and Gemini), each receiving an identical standardized prompt that embeds the same rubric, and by faculty serving as the reference standard.
To avoid overloading faculty, full double human grading may not be feasible; the human reference will therefore consist of expert faculty grading by one independent rater or, when resources allow, two independent raters. In contrast, paired assessment is fully implemented across the AI models: each examination is scored by several LLMs, and each model is queried in duplicate, allowing the study to estimate agreement between models and the test-retest stability of each model.
The primary aim is to quantify agreement between LLM-generated scores and the faculty reference score. Secondary aims include agreement among the LLMs, test-retest reliability of each model, criterion-level agreement, the quality and usefulness of the qualitative feedback generated, the time and cost associated with each approach, and students' perceptions of the usefulness of human versus AI feedback.
The findings will clarify the strengths and limitations of LLMs as supportive tools for formative assessment in health-professions education and will inform criteria for their responsible and effective use. No LLM output will affect students' official grades, which remain the sole responsibility of faculty.
详细描述
BACKGROUND AND RATIONALE The assessment of clinical case examinations in physiotherapy requires the appraisal of multiple dimensions, including clinical reasoning, selection of appropriate techniques, treatment dosage, ethical considerations, and patient communication. This grading process demands substantial faculty time and may be affected by evaluator fatigue and inter-rater variability. Large language models (LLMs) have shown notable capabilities in text comprehension, reasoning, and the generation of structured feedback, and preliminary evidence suggests they may provide consistent evaluations in medical education contexts. However, questions remain regarding their reliability, potential bias, and their ability to capture the complexity of clinical reasoning. There is currently limited empirical evidence on whether LLMs can complement human grading while maintaining quality standards and offering immediate formative feedback. This study addresses that gap through a systematic comparison between human assessment (reference standard) and LLM-assisted assessment using identical materials and criteria.
OBJECTIVES Primary objective: To quantify the agreement between the scores generated by LLMs and the faculty reference score in the assessment of physiotherapy clinical-reasoning case examinations.
Secondary objectives:
- To estimate the agreement among different LLMs (inter-model reliability).
- To estimate the test-retest reliability of each LLM (intra-model reliability) when the same examination is scored on repeated, independent administrations.
- To evaluate agreement at the level of individual rubric criteria.
- To compare the quality, specificity, and formative usefulness of the qualitative feedback produced by LLMs and by faculty.
- To compare the time and cost associated with human and LLM assessment.
- To assess students' perceptions of the usefulness, fairness, and transparency of human versus AI feedback.
STUDY DESIGN Cross-sectional inter-rater agreement and reliability study with repeated measures, in which the same set of de-identified student clinical case examinations is assessed independently by human and artificial-intelligence raters using a shared, predefined rubric. The study is observational and educational in nature; it does not modify the teaching or the official assessment received by students.
研究设计
- 研究类型
- Observational
- 观察模型
- Cohort
- 时间视角
- Cross Sectional
入排标准
- 年龄范围
- 18 Years 至 —(Adult, Older Adult)
- 性别
- All
- 接受健康志愿者
- 是
入选标准
- •Students officially enrolled in the course "Specific Methods in Physiotherapy" (third year of the Physiotherapy Degree) during the study period.
- •Submission of a completed written clinical-reasoning case examination as part of the course.
- •Provision of informed consent for the anonymized examination to be used for educational-research purposes.
排除标准
- •Refusal to provide, or withdrawal of, informed consent.
- •Blank, incomplete, or non-evaluable examinations (e.g., no developed written response).
- •Examinations that cannot be reliably de-identified prior to assessment.
结局指标
主要结局
Agreement between LLM global scores and the faculty reference global score
时间窗: Single cross-sectional assessment during the data-collection period (approximately 2 months)
Agreement between the global examination score generated by each large language model (LLM) and the faculty reference global score, computed for the same anonymized examinations. Agreement is quantified with the intraclass correlation coefficient (ICC), two-way random-effects model, absolute-agreement definition, single- and average-measures forms \[ICC(2,1) and ICC(2,k)\], with 95% confidence intervals. Systematic bias is examined with Bland-Altman analysis (mean difference and 95% limits of agreement). ICC is interpreted as poor (\<0.50), moderate (0.50-0.75), good (0.75-0.90), or excellent (\>0.90). Pre-specified target: ICC \>= 0.75.
次要结局
- Criterion-level agreement between LLM and faculty scores(Single cross-sectional assessment during the data-collection period (approximately 2 months))
- Intra-model test-retest reliability of each large language model(Single cross-sectional assessment during the data-collection period (approximately 2 months))
- Quality and coverage of the qualitative feedback(Assessed after completion of all evaluations, during the analysis period (approximately 3 months))
- Mean evaluation time per examination: faculty versus LLM(Single cross-sectional assessment during the data-collection period (approximately 2 months))
- Cost per evaluation: faculty versus LLM(Single cross-sectional assessment during the data-collection period (approximately 2 months))
研究者
Alfredo Lerín Calvo
Mr.
Neuron, Spain
