跳至主要内容
临床试验/NCT07458971
NCT07458971招募中不适用

ICF-Based Biopsychosocial Assessment With Artificial Intelligence-Assisted Profile Prediction: Trapeziometacarpal Osteoarthritis Model

Hacettepe University2 个研究点 分布在 1 个国家目标入组 93 人开始时间: 2026年2月19日最近更新:

试验速览

阶段
不适用
状态
招募中
入组人数
93
试验地点
2
主要终点
Large Language Model Clinical Knowledge Accuracy

研究概览

简要总结

Trapeziometacarpal osteoarthritis (TMC OA) is a common condition affecting the base of the thumb that causes pain, weakness, and difficulty with daily hand use. Current clinical assessment often focuses on physical findings alone, without considering psychological and social factors that also influence patient outcomes.

This study has three objectives organized as interrelated work packages:

OBJECTIVE 1 (Clinical Assessment): To comprehensively assess individuals with TMC OA using the International Classification of Functioning, Disability and Health (ICF) framework. This includes evaluating pain, joint mobility, grip strength, daily activity limitations, social participation, psychological factors (anxiety, depression, fear of movement, pain beliefs), and environmental factors (family support, ergonomic adaptations).

OBJECTIVE 2 (AI Knowledge Evaluation): To compare the performance of four large language models (GPT-5.2, Claude Opus 4.6, Gemini 2.5 Pro, LLaMA 4 Maverick) in answering patient questions about TMC OA, and to test whether adding a clinician-oriented system prompt changes that performance. Each question is submitted under two prompting conditions (zero-shot and clinician-oriented system prompt) and responses are rated by blinded experts for accuracy, comprehensiveness and clinical relevance, with readability assessed by automated indices.

OBJECTIVE 3 (AI-Based Prediction): To analyze whether the best-performing large language model can predict multidimensional ICF-based patient profiles using only a limited set of core clinical parameters.

详细描述

This research consists of three independent but interrelated work packages with different methods and targets.

Work Package 1 (Clinical Data Collection and ICF-Based Profile Analysis): Participants with TMC OA will undergo a single face-to-face comprehensive assessment using a cross-sectional design. The assessment battery is structured according to the ICF framework and covers five domains: (a) Body Structure/Function: pain, joint mobility, grip and pinch strength, joint stability, and OA staging; (b) Activity: daily activity limitations and pain-activity patterns (avoidance, overdoing, pacing); (c) Participation: social, domestic, and occupational participation; (d) Personal Factors: pain beliefs, coping strategies, kinesiophobia, anxiety, and depression; (e) Environmental Factors: family support and ergonomic adaptations.

Work Package 2 (Comparison of Large Language Models' Clinical Knowledge Performance): Forty patient questions derived from search-engine "People Also Ask" data are submitted programmatically, through provider application programming interfaces, to four large language models (GPT-5.2, Claude Opus 4.6, Gemini 2.5 Pro, LLaMA 4 Maverick) under two prompting conditions: zero-shot, and with a clinician-oriented system prompt. This yields 320 responses. Two experts (a hand therapist and a hand surgeon), blinded to model identity and prompting condition, independently rate each response for accuracy, comprehensiveness and clinical relevance on anchored 5-point scales. Criterion-level disagreements of two points or more are referred to a third blinded hand surgeon, whose rating is substituted for that criterion. Readability is computed automatically as Flesch Reading Ease and Flesch-Kincaid Grade Level.

Work Package 3 (LLM-Based Predictive Profile Modeling): The best-performing LLM identified in WP2 will be provided with core clinical predictors from WP1 data. The model's predictions for multidimensional ICF-based patient profiles will be compared against actual assessment results using established agreement and performance metrics.

Sample size: Based on a priori power analysis (alpha=0.05, power=0.80, effect size=0.131), a minimum of 93 participants is required.

研究设计

研究类型
Observational
观察模型
Cohort
时间视角
Cross Sectional

入排标准

年龄范围
25 Years 至 74 Years(Adult, Older Adult)
性别
All
接受健康志愿者

入选标准

  • Diagnosis of trapeziometacarpal osteoarthritis (TMC OA) confirmed by an orthopedic surgeon and/or hand surgeon
  • TMC OA-related symptoms persisting for more than 3 months
  • Aged between 25 and 74 years
  • Literate in Turkish
  • Adequate cognitive function (Mini-Mental State Examination score of 25 or above)
  • No other chronic systemic disease (e.g., rheumatoid arthritis, chronic diabetes, cardiovascular disease, chronic hepatitis)
  • Voluntary participation with signed informed consent

排除标准

  • Unwillingness to participate
  • Presence of a different orthopedic condition or prior surgery involving the thumb on the unilateral upper extremity
  • Uncontrolled systemic diseases (chronic obstructive pulmonary disease, congestive heart failure, endocrine system disease, history of stroke)
  • Diagnosis of any major psychopathology and currently receiving psychiatric or psychological treatment
  • Pregnancy

研究组 & 干预措施

TMC OA Group

Individuals diagnosed with trapeziometacarpal osteoarthritis (TMC OA) by an orthopedic surgeon and/or hand surgeon, with symptoms persisting for more than 3 months, aged 25-74 years, referred to the Hand Surgery Rehabilitation Unit. All participants will undergo the same ICF-based comprehensive biopsychosocial assessment battery in a single face-to-face session.

结局指标

主要结局

Large Language Model Clinical Knowledge Accuracy

时间窗: Baseline (single assessment during the data collection period)

Accuracy of responses from four large language models to 40 patient-derived questions, each submitted under two prompting conditions (320 responses). Two blinded experts rate each response on an anchored 5-point scale (1 = mostly incorrect or potentially harmful, 5 = no factual error); criterion-level disagreements of two points or more are resolved by a third blinded expert. Higher scores indicate greater accuracy.

Large Language Model Content Comprehensiveness

时间窗: Baseline (single assessment at enrollment)

Coverage of the elements a specialist would expect in an answer, rated by two blinded experts on an anchored 5-point scale (1 = question not genuinely answered, 5 = all essential elements covered), with third-expert arbitration of disagreements of two points or more.

Large Language Model Clinical Relevance

时间窗: Baseline (single assessment)

Alignment of the response with the question asked, rated on an anchored 5-point scale (1 = unrelated, 5 = precisely and fully answers the question), with third-expert arbitration.

Large Language Model Readability Score

时间窗: Baseline (calculated immediately after response generation)

Readability of generated responses, computed automatically as the Flesch Reading Ease score and the Flesch-Kincaid Grade Level. Higher Flesch Reading Ease scores indicate easier text; higher Flesch-Kincaid values indicate a higher required reading grade.

Grip Strength

时间窗: Baseline (single assessment at enrollment)

Measured using a Jamar dynamometer. The participant performs the test in a seated position with the elbow flexed at 90 degrees. Unit of Measure: Kilograms

Pinch Strength

时间窗: Baseline (single assessment at enrollment)

Measured using a pinchmeter to assess tip-to-tip and key pinch strength. Unit of Measure: Kilograms

Thumb Opposition (Kapandji Score)

时间窗: Baseline (single assessment at enrollment)

Assessment of thumb opposition using the Kapandji score, which ranges from 0 to 10. Higher scores indicate better thumb opposition and mobility.

Pain Intensity

时间窗: Baseline (single assessment at enrollment)

Measured using a Visual Analog Scale (VAS) ranging from 0 (no pain) to 10 (worst imaginable pain). Higher scores indicate greater pain intensity.

Pain Duration

时间窗: Baseline (single assessment at enrollment)

Total duration of thumb pain reported by the participant.

Radiographic Severity (Eaton-Littler Stage)

时间窗: Baseline (single assessment at enrollment)

Evaluation of the trapeziometacarpal joint osteoarthritis stage based on the Eaton-Littler classification (Stages I through IV).

Radial Subluxation Ratio

时间窗: Baseline (single assessment at enrollment)

Radiographic measurement of the radial subluxation of the metacarpal base on the trapezium.

Upper Extremity Disability (QuickDASH)

时间窗: Baseline (single assessment at enrollment)

Measured using the Quick Disabilities of the Arm, Shoulder and Hand (QuickDASH) questionnaire. The score ranges from 0 to 100, where higher scores indicate greater disability and symptoms.

Hand Disability (Turkish Thumb Disability Index - TDX)

时间窗: Baseline (single assessment at enrollment)

Assessment of thumb-related disability. Scores range from 0 to 100, with higher scores indicating greater functional impairment.

Joint Hypermobility (Beighton Score)

时间窗: Baseline (single assessment at enrollment)

Assessment of generalized joint laxity using the Beighton score. The total score ranges from 0 to 9, where higher scores indicate greater hypermobility.

Thumb Joint Range of Motion

时间窗: Baseline (single assessment at enrollment)

Active range of motion of the thumb joints measured using a goniometer. Unit of Measure: Degrees

Provocative Tests

时间窗: Baseline (single assessment at enrollment)

Clinical assessment using metacarpal adduction and extension tests to provoke symptoms. Presence or absence of pain (Binary: Yes/No)

Environmental Factors: Social Support and Ergonomic Adaptations

时间窗: Baseline (single assessment at enrollment)

Qualitative assessment of the participant's family support and the presence of ergonomic adaptations in their daily environment.

Emotional Status (Hospital Anxiety and Depression Scale)

时间窗: Baseline (single assessment at enrollment)

Measured using the Hospital Anxiety and Depression Scale (HADS), which consists of two subscales: Anxiety (HADS-A) and Depression (HADS-D). Each subscale ranges from 0 to 21, where higher scores indicate greater levels of anxiety or depression (worse outcome).

Kinesiophobia Level (Tampa Scale of Kinesiophobia)

时间窗: Baseline (single assessment at enrollment)

Measured using the 17-item Tampa Scale of Kinesiophobia (TSK-17) to assess the fear of movement or re-injury. Total scores range from 17 to 68, where higher scores indicate greater kinesiophobia (worse outcome).

Pain-Activity Patterns (Patterns of Activity Measure-Pain).

时间窗: Baseline (single assessment at enrollment).

Measured using the Patterns of Activity Measure-Pain (POAM-P) questionnaire to classify participants into three patterns: avoidance, overdoing, and pacing. Each subscale score indicates the frequency of that specific activity pattern. Higher scores on each subscale indicate a more frequent use of that specific activity pattern.

Pain Beliefs Profile (Pain Beliefs Questionnaire)

时间窗: Baseline (single assessment at enrollment).

Assessed using the Pain Beliefs Questionnaire (PBQ), which evaluates two dimensions: Organic and Psychological pain beliefs. Scores range from 1 to 6 for each subscale, where higher scores indicate a stronger belief in that specific dimension (e.g., higher organic scores mean a stronger belief that pain is due to physical damage).

Pain Coping Strategies (Pain Coping Questionnaire).

时间窗: Baseline (single assessment at enrollment).

Measured using the Pain Coping Questionnaire (PCQ) to assess the frequency of different coping strategies (e.g., information seeking, problem solving, distraction). Higher scores indicate a more frequent use of the respective coping strategy.

Large Language Model Response Reproducibility

时间窗: Within 24 hours after initial query

Assessment of content consistency between two repeated queries of the same questions. Evaluated based on the percentage of agreement between the two sets of responses.

LLM Prediction Accuracy for Continuous ICF Profiles

时间窗: Within 3 months after the completion of clinical data collection.

Prediction accuracy of the best-performing LLM (identified in WP2) in estimating continuous clinical scores (e.g., Grip Strength, QuickDASH scores) from core clinical predictors. Accuracy will be measured using the Intraclass Correlation Coefficient (ICC) to evaluate the agreement between LLM-predicted values and actual clinical assessment results.

LLM Prediction Accuracy for Categorical ICF Profiles

时间窗: Within 3 months after the completion of clinical data collection.

Prediction accuracy of the best-performing LLM in estimating categorical patient profiles (e.g., Eaton-Littler Stage, POAM-P activity patterns). Accuracy will be measured using Cohen's Kappa coefficient to evaluate the agreement between LLM-predicted categories and actual expert-diagnosed categories.

次要结局

  • Correlations Between Pain-Activity Patterns and Clinical Variables(Baseline (single assessment at enrollment))
  • Kinesiophobia Level (Tampa Kinesiophobia Scale)(Baseline (single assessment at enrollment).)
  • Pain Beliefs Profile (Pain Beliefs Questionnaire)(Baseline (single assessment at enrollment))
  • Anxiety and Depression (Hospital Anxiety and Depression Scale)(Baseline (single assessment at enrollment).)
  • Pain Coping Strategies (Pain Coping Questionnaire)(Baseline (single assessment at enrollment))

研究者

申办方类型
Other
责任方
Principal Investigator
主要研究者

Cigdem Ayhan

Professor

Hacettepe University

研究点 (2)

Loading locations...

相似试验