跳至主要内容
临床试验/NCT07824531
NCT07824531进行中(未招募)不适用

Feasibility of a Self-Hosted Vision-Language Model for Detecting Caregiver Contact and External Support, the Observable Criteria That Define GMFCS Levels, From Multi-View Movement Video in Children Aged 2 to 6 Years With Cerebral Palsy: A Two-Center Diagnostic Accuracy Study

Samsung Medical Center2 个研究点 分布在 1 个国家目标入组 26 人开始时间: 2025年5月30日最近更新:
适应症
干预措施

试验速览

阶段
不适用
状态
进行中(未招募)
入组人数
26
试验地点
2
主要终点
Accuracy of the model-determined adult-contact call against the human annotator's record

研究概览

简要总结

The GMFCS sorts children with cerebral palsy into five levels of gross motor function. What mainly separates one level from the next is two things an observer can see: whether another person has to physically help the child move, and whether the child has to bear weight on something external, such as a walker, a support stand, or furniture. Scoring these takes an experienced clinician, and it is hardest between the ages of 2 and 6.

This study asks a narrow question. Can a vision-language model, meaning an artificial-intelligence model that looks at images and answers questions about them, detect those two things from short video clips of a child moving?

Twenty-six children with cerebral palsy aged 2 to 6 years were filmed at two hospitals. Each child performed everyday movements, including walking, crawling, rolling sideways, rising from the floor to standing, and lowering back to the floor, and up to three cameras recorded every attempt at the same time. Sixteen still frames were taken from each recorded attempt, always sixteen, and shown to the model without the child's identity, the name of the movement, or the child's GMFCS level.

The model answers two yes-or-no questions and nothing else: did an adult touch the child, and did the child bear weight on an external object. A fixed rule, not the model, turns those answers into one determination per attempt, namely whether the child performed the movement unaided. These determinations are compared with what a human annotator recorded for the same clips while blind to clinical information.

All rating is done with open-weight models running on hardware the investigators control, and every reported figure is tied to the specific model version that produced it.

The study does not assign a GMFCS level, and it is not built to. It tests whether the two observations the GMFCS itself relies on can be read from video reliably. If they can, they can be placed in front of a clinician as evidence, which is where later work on supporting GMFCS assessment would start.

详细描述

BACKGROUND

The GMFCS-E&R separates its five levels mainly by two things: what a child cannot do without help from another person, and what a child cannot do without a hand-held mobility device. The wording is explicit in the age bands relevant here. At Level I a child moves in and out of floor sitting and standing "without adult assistance." At Level III a child "may require adult assistance to assume sitting" and needs "adult assistance for steering and turning" when walking with a walker. Adult contact and external support are therefore not proxies chosen for convenience. They are part of the definition.

Scoring them is another matter. It takes an experienced clinician, it is slow, and between 2 and 6 years of age the assignment is known to be difficult.

STUDY DESIGN

This is a two-center feasibility study of a determination made from video. The index test is a binary determination produced by an open-weight vision-language model. The reference standard is a human annotator's record of the same clips. No clinical grade is requested from the model at any point.

研究设计

研究类型
Observational
观察模型
Cohort
时间视角
Cross Sectional

入排标准

年龄范围
2 Years 至 6 Years(Child)
性别
All
接受健康志愿者

入选标准

  • Clinical diagnosis of cerebral palsy
  • Aged 2 to 6 years at the time of video recording
  • Under the care of the department of physical and rehabilitation medicine at Samsung Medical Center or Asan Medical Center
  • A GMFCS level assigned by the treating clinician
  • Written informed consent from a parent or legal guardian for video recording

排除标准

  • Consent for video recording is not given
  • GMFCS level cannot be clearly assessed
  • The child or the parent or legal guardian declines to take part in the study
  • Withdrawal Criteria:
  • The child or the parent or legal guardian requests withdrawal during the study
  • The collected video is not of adequate quality for analysis
  • The participant's health status changes substantially during the study period

研究组 & 干预措施

Standardized Video Assessment

All 26 enrolled children. Each child performed five prescribed gross motor tasks (walking, crawling, side-rolling, floor sit-to-stand, stand-to-floor-sit) while being recorded simultaneously by up to three cameras. Every recorded attempt was submitted to the index test, in which 16 extracted frames are read by an open-weight vision-language model. There is no comparison group.

干预措施: Vision-language model assessment of movement video (Diagnostic Test)

结局指标

主要结局

Accuracy of the model-determined adult-contact call against the human annotator's record

时间窗: Single video assessment per participant; recordings collected May 2025 to June 2026

For each movement attempt in the fixed 178-attempt floor-transition set, the rater model's binary determination of whether an adult touched the child is compared with the human annotator's blind record for the same attempt. Accuracy is the proportion of attempts in agreement, reported as a count out of 178. Sensitivity and specificity against the same reference are reported alongside. Every rater model is evaluated on identical image inputs, so accuracies are directly comparable, and comparisons between rater models use McNemar's paired test. The human record is the sole reference.

Accuracy of the model-determined external-support call against the human annotator's record

时间窗: Single video assessment per participant; recordings collected May 2025 to June 2026

For each movement attempt in the same fixed 178-attempt set, the rater model's binary determination of whether the child bore load through an external object is compared with the human annotator's blind record of acrylic support stand or walker use. Accuracy is reported as a count out of 178, with sensitivity and specificity against the same reference.

次要结局

  • Stratified concordance between model and human contact determinations within participant and movement strata(Single video assessment per participant; recordings collected May 2025 to June 2026)
  • Agreement between independent rater models on identical image inputs(Single video assessment per participant; recordings collected May 2025 to June 2026)
  • Accuracy of image-independent baselines on the same attempts(Single video assessment per participant; recordings collected May 2025 to June 2026)

研究者

申办方类型
Other
责任方
Principal Investigator
主要研究者

Jeong Yi Kwon

Professor, Department of Physical and Rehabilitation Medicine

Samsung Medical Center

研究点 (2)

Loading locations...

相似试验