Effect of the SARA Citation-Verification Framework on Medical Students' Verification of AI-Generated Citation Hallucinations: A Randomized Controlled Trial With an Explanatory Sequential Mixed-Methods Component
试验速览
- 阶段
- 不适用
- 状态
- 尚未招募
- 发起方
- 入组人数
- 320
- 主要终点
- Independent Citation-Verification Performance on an Unstructured Scholarly Transfer Task
研究概览
简要总结
Generative artificial intelligence (GenAI) tools are increasingly used by medical students for literature searching, scientific writing, and research. However, AI-generated citations may contain errors. Some references may not exist or may contain incorrect bibliographic details, while other references may be genuine but may not actually support the claim for which they are cited.
This randomized educational study will evaluate whether a structured citation-verification framework called SARA can improve medical students' ability to identify and correct these problems. SARA teaches students to assess four elements of a citation: Source authenticity, Accuracy, Relevance, and Alignment between the cited evidence and the claim being made.
Year 3-5 medical students will be randomly assigned to either a 75-minute SARA training session or an attention-matched session on responsible use of GenAI in academic work. Participants will complete citation-verification assessments before training, immediately after training, and 12 weeks later. The main outcome will assess whether students can independently apply the verification approach to an unstructured literature-review task 12 weeks after training, without being prompted to use SARA.
The study will also examine which types of citation errors remain difficult, whether students initiate appropriate verification, the quality of their verification, handling of valid citations, and confidence in their judgments. A subgroup of participants will take part in think-aloud and follow-up interviews to explore how they approach citation verification and transfer the strategy to new tasks.
详细描述
AI-generated citation errors can occur at different levels. Some are source-level errors, such as fabricated references or incorrect bibliographic details. Others are more difficult to detect because the cited publication is genuine, but the evidence may be irrelevant, may not support the direction, magnitude, population, or conclusion of the claim, or may be presented without important qualifiers or limitations. In this study, these problems are examined across a five-level citation-error taxonomy ranging from non-existent sources to contextual omission or overgeneralization.
The SARA framework was developed as an educational scaffold to make citation-verification reasoning explicit. It prompts learners to ask whether the source exists, whether the citation details are accurate, whether the source is relevant to the specific claim, and whether the evidence is aligned with the claim as written. The intervention uses modelling, coached practice, feedback, articulation, reflection, and progressive removal of support so that learners move from guided verification toward independent judgment.
A key feature of the study is assessment of delayed transfer rather than immediate post-training performance alone. At 12 weeks, participants will review a realistic literature-review passage containing both valid and problematic citation-claim relationships without being told how many errors are present or being prompted to use SARA. This is intended to evaluate whether the verification strategy has become sufficiently internalized to be applied independently in a new scholarly context.
The explanatory qualitative component will explore why some learners successfully transfer verification skills while others continue to miss deeper evidence-claim problems. Think-aloud and stimulated-recall interviews will examine what triggers verification, how learners decide that evidence is sufficient, when they stop checking, and how they distinguish source existence from evidentiary support.
研究设计
- 研究类型
- 干预性
- 分配方式
- 随机
- 干预模型
- 平行分组
- 主要目的
- 其他
- 盲法
- 单盲 (结局评估者)
入排标准
- 性别
- All
- 接受健康志愿者
- 是
入选标准
- Enrolled as an MBBS student in Year 3, Year 4, or Year 5 at Quaid-e-Azam Medical College, Bahawalpur.
- Has received basic curricular exposure to research methodology or scientific literature searching.
- Has not previously received formal structured training specifically in verification of AI-generated citations.
- Willing and able to provide written informed consent.
- Meets the age eligibility specified in the approved study protocol.
排除标准
- Previous participation in development, cognitive interviewing, piloting, or validation testing of the SARA assessment materials.
- Previous formal training in the SARA citation-verification framework before enrolment.
- Unwilling or unable to provide informed consent.
研究组 & 干预措施
SARA Citation-Verification Training
Participants assigned to this arm will receive a standardized 75-minute educational intervention using the SARA citation-verification framework: Source authenticity, Accuracy, Relevance, and Alignment. Training includes facilitator modelling and think-aloud demonstration, coached practice with feedback, use of a one-page verification scaffold, articulation and reflection, and progressive fading of prompts before independent practice. The intervention is designed to teach students to distinguish source-level citation errors from deeper evidence-claim errors and to apply structured verification independently.
干预措施: SARA Citation-Verification Training (Behavioral)
Responsible GenAI Education Control
Participants assigned to this arm will receive an attention-matched 75-minute educational session on responsible use of generative artificial intelligence in academic work. The session covers privacy, disclosure, authorship, academic integrity, and general awareness that GenAI may produce errors. It does not teach the SARA framework or another structured citation-verification process. The session is matched to the intervention for duration, group size, facilitator interaction, and delivery format.
干预措施: Responsible GenAI Education Intervention Description (Behavioral)
结局指标
主要结局
Independent Citation-Verification Performance on an Unstructured Scholarly Transfer Task
时间窗: 12 weeks after the intervention
Participants will review a novel approximately 650-word AI-generated literature-review passage containing eight embedded citation-claim relationships. The passage includes both valid and erroneous citations spanning the study's five-level citation-error taxonomy. Participants will not be prompted to use the SARA framework and will not be told how many or what types of errors are present. Each citation-claim relationship will be scored against a prespecified gold standard for the appropriateness of the participant's final disposition. One point will be awarded for each appropriate disposition, producing a total score from 0 to 8, with higher scores indicating better independent citation-verification performance. Responses will be scored by outcome assessors blinded to intervention allocation.
次要结局
- Structured Citation-Verification Performance(Immediately after the intervention and 12 weeks after the intervention)
- Citation-Verification Accuracy by Citation-Error Condition(Immediately after the intervention and 12 weeks after the intervention)
- Verification Initiation(Immediately after the intervention and 12 weeks after the intervention)
- High-Quality Citation Verification(12 weeks after the intervention)
- Unnecessary Alteration or Rejection of Valid Citations(12 weeks after the intervention)
- Confidence Calibration in Citation-Verification Judgments(12 weeks after the intervention)
- Retention of Structured Citation-Verification Performance(Baseline, immediately after the intervention, and 12 weeks after the intervention)
研究者
Sara Reza
Associate Professor
Quaid-e-Azam Medical College
标识符
- NCT 编号
- NCT07844746
- 其他研究编号
- QAMC-SARA-2026-01
日期
- 首次提交
- (9天前)
- 首次发布
- (前天)
- 主要完成日期
- (3个月后)
- 研究完成日期
- (4个月后)
- 最近核实
- (29天前)
- 最近更新
- (前天)
监管与共享
- FDA 监管药物
- 否
- FDA 监管器械
- 否
- 个体参与者数据共享计划
- 是
- 是否有结果
- 否
De-identified individual participant quantitative data underlying the results reported in the primary publication will be made available. A data dictionary and relevant supporting documentation will also be provided. Qualitative interview transcripts will not be shared publicly because of the risk of participant re-identification.
