跳至主要内容
临床试验/NCT07817238
NCT07817238进行中(未招募)不适用

AI-Assisted Electrode Contact Configuration Mapping for Epidural Electrical Stimulation in Spinal Cord Injury: A Comparative Evaluation of Large Language Models

Istanbul Gelisim University1 个研究点 分布在 1 个国家目标入组 20 人开始时间: 2026年9月2日最近更新:
适应症
干预措施

试验速览

阶段
不适用
状态
进行中(未招募)
发起方
入组人数
20
试验地点
1
主要终点
Clinical Accuracy Score of Large Language Model Responses

研究概览

简要总结

This observational and methodological study aims to compare the performance of large language models in generating electrode contact configuration recommendations for epidural electrical stimulation in spinal cord injury.

Five standardized synthetic spinal cord injury scenarios will be presented to four large language models: ChatGPT-4o, Claude, Grok 3, and Gemini 2.5 Pro. Each model will receive the same standardized prompt. The generated responses will be anonymized and evaluated independently by experts with experience in spinal cord injury rehabilitation and epidural electrical stimulation.

The responses will be assessed in five main areas: clinical accuracy, technical feasibility, safety awareness, consistency with current clinical guidance, and completeness of the response. Agreement between expert evaluators will also be examined.

No real patients, human participants, clinical interventions, or personal health data are included in this study. The study is designed to explore the potential and current limitations of large language models as artificial intelligence-based clinical decision-support tools in neurorehabilitation.

研究设计

研究类型
Observational
观察模型
Other
时间视角
Cross Sectional

入排标准

性别
All
接受健康志愿者

入选标准

  • Responses generated for one of the five predefined standardized synthetic spinal cord injury scenarios.
  • Responses generated using the identical standardized prompt specified in the study protocol.
  • Responses generated by one of the four prespecified large language models.
  • Complete responses available for expert evaluation.

排除标准

  • Responses generated using prompts that differ from the standardized study prompt.
  • Incomplete, interrupted, or technically corrupted model outputs.
  • Duplicate responses or outputs not corresponding to a predefined synthetic scenario.
  • Any response generated using real patient-identifiable or personal health information.

研究组 & 干预措施

Gemini 2.5 Pro

Responses generated by Grok 3 for five standardized synthetic spinal cord injury scenarios using the same standardized prompt. The responses will be evaluated for clinical accuracy, technical feasibility, safety awareness, guideline consistency, and completeness.

干预措施: Gemini 2.5 Pro Large Language Model (Other)

Claude

Responses generated by Claude for five standardized synthetic spinal cord injury scenarios using the same standardized prompt. The responses will be evaluated for clinical accuracy, technical feasibility, safety awareness, guideline consistency, and completeness.

干预措施: Claude Large Language Model (Other)

Grok 3

Responses generated by Grok 3 for five standardized synthetic spinal cord injury scenarios using the same standardized prompt. The responses will be evaluated for clinical accuracy, technical feasibility, safety awareness, guideline consistency, and completeness.

干预措施: Grok 3 Large Language Model (Other)

ChatGPT-4o

Responses generated by ChatGPT-4o for five standardized synthetic spinal cord injury scenarios using the same standardized prompt. The responses will be evaluated for clinical accuracy, technical feasibility, safety awareness, guideline consistency, and completeness.

干预措施: ChatGPT-4o Large Language Model (Other)

结局指标

主要结局

Clinical Accuracy Score of Large Language Model Responses

时间窗: At the time of expert evaluation, within 1 week after study initiation

Clinical accuracy of the epidural electrical stimulation electrode contact configuration recommendations generated by each large language model will be independently evaluated by expert reviewers using a 5-point Likert-type rating scale. Higher scores indicate greater clinical accuracy of the generated recommendations.

次要结局

  • Technical Feasibility Score of Large Language Model Responses(At expert evaluation, within 1 week after study initiation)
  • Safety Awareness Score of Large Language Model Responses(At expert evaluation, within 1 week after study initiation)
  • Clinical Guideline Consistency Score of Large Language Model Responses(At expert evaluation, within 1 week after study initiation)
  • Response Completeness Score of Large Language Model Responses(At expert evaluation, within 1 week after study initiation)

研究者

发起方
Istanbul Gelisim University
申办方类型
Other
责任方
Principal Investigator
主要研究者

Görkem Açar

Director

Istanbul Gelisim University

研究点 (1)

Loading locations...

相似试验

Large Language Models for Epidural Stimulation... | 临床试验