跳至主要内容
临床试验/NCT07835347
NCT07835347已完成不适用

Alignment of Large Language Model Responses on Stuttering Definition, Assessment, and Therapy With Clinical Practice Guidelines: A Multi-Model Comparison

Istanbul Gelisim University1 个研究点 分布在 1 个国家目标入组 5 人开始时间: 2026年8月7日最近更新:
适应症
干预措施

试验速览

阶段
不适用
状态
已完成
发起方
入组人数
5
试验地点
1
主要终点
Expert-Rated Quality of Large Language Model Responses

研究概览

简要总结

This observational study aims to compare the quality of responses generated by four large language models (ChatGPT, Claude, Google Gemini, and Grok-4.20) to questions about stuttering. The study focuses on three main areas: general information about stuttering, clinical assessment, and therapy. A total of nine questions were developed based on evidence-based clinical practice guidance, including the American Speech-Language-Hearing Association (ASHA) Practice Portal. Each question is presented to each language model in separate sessions, and the generated responses are recorded for evaluation.

Five speech-language therapists with clinical experience in stuttering independently evaluate the model-generated responses. Each response is rated for relevance, accuracy, clarity, completeness, and consistency using a 5-point Likert scale. The responses are also compared with guideline-based reference information. The study does not involve any clinical intervention or treatment of patients. The aim is to determine how closely large language model responses align with current clinical practice guidance and to identify their potential strengths and limitations when used as informational or clinical support tools in the field of stuttering.

研究设计

研究类型
Observational
观察模型
Other
时间视角
Cross Sectional

入排标准

年龄范围
18 Years 至 —(Adult, Older Adult)
性别
All
接受健康志愿者
是

入选标准

  • •Having an undergraduate or postgraduate degree in speech and language therapy.
  • •Actively practicing clinically in the field of stuttering/fluency disorders.
  • •Voluntarily agreeing to participate in the study.
  • •Completing the evaluation form in full.

排除标准

  • •Not having active clinical experience in the field of stuttering.
  • •Incomplete completion of the evaluation form.
  • •Withdrawal of voluntary participation during the evaluation process.

研究组 & 干预措施

Expert Speech-Language Therapists

Five speech-language therapists with clinical experience in stuttering independently evaluated responses generated by four large language models (ChatGPT, Claude, Google Gemini, and Grok-4.20). The responses addressed nine questions covering general information about stuttering, clinical assessment, and therapy. Each response was rated for relevance, accuracy, clarity, completeness, and consistency using a 5-point Likert scale.

干预措施: ChatGPT-Generated Responses (Other)

Expert Speech-Language Therapists

Five speech-language therapists with clinical experience in stuttering independently evaluated responses generated by four large language models (ChatGPT, Claude, Google Gemini, and Grok-4.20). The responses addressed nine questions covering general information about stuttering, clinical assessment, and therapy. Each response was rated for relevance, accuracy, clarity, completeness, and consistency using a 5-point Likert scale.

干预措施: Claude-Generated Responses (Other)

Expert Speech-Language Therapists

Five speech-language therapists with clinical experience in stuttering independently evaluated responses generated by four large language models (ChatGPT, Claude, Google Gemini, and Grok-4.20). The responses addressed nine questions covering general information about stuttering, clinical assessment, and therapy. Each response was rated for relevance, accuracy, clarity, completeness, and consistency using a 5-point Likert scale.

干预措施: Google Gemini-Generated Responses (Other)

Expert Speech-Language Therapists

Five speech-language therapists with clinical experience in stuttering independently evaluated responses generated by four large language models (ChatGPT, Claude, Google Gemini, and Grok-4.20). The responses addressed nine questions covering general information about stuttering, clinical assessment, and therapy. Each response was rated for relevance, accuracy, clarity, completeness, and consistency using a 5-point Likert scale.

干预措施: Grok-4.20-Generated Responses (Other)

结局指标

主要结局

Expert-Rated Quality of Large Language Model Responses

时间窗: During the single cross-sectional expert evaluation conducted over approximately 1 month

Responses generated by four large language models (ChatGPT, Claude, Google Gemini, and Grok-4.20) to nine standardized questions about stuttering were independently evaluated by five speech-language therapists with clinical experience in stuttering. Each response was rated using a 5-point Likert scale (1 = strongly disagree to 5 = strongly agree) across five criteria: relevance, accuracy, clarity, completeness, and consistency. Responses were evaluated against guideline-based reference information derived from evidence-based clinical practice guidance. Higher scores indicate better response quality and closer alignment with clinical practice guidance. Mean scores were calculated for each large language model and each evaluation criterion.

次要结局

未报告次要终点

研究者

发起方
Istanbul Gelisim University
申办方类型
Other
责任方
Principal Investigator
主要研究者

Esra EROL

Director

Istanbul Gelisim University

研究点 (1)

Loading locations...

相似试验

Large Language Models for Stuttering Assessment and... | 临床试验