Development and Validation of Multimodal Deep Learning Model for Autonomous Diagnosis, Generative Reporting, and Specialist Referral in Ophthalmic Diseases: An International Multicenter Cohort Study
Trial Snapshot
- Phase
- Not Applicable
- Status
- Recruiting
- Enrollment
- 2,000
- Locations
- 1
- Primary Endpoint
- Diagnostic accuracy of multimodal vision-language model.
Study Overview
Brief Summary
Accurate and comprehensive interpretation of anterior segment diseases from slit-lamp and smartphone photographs remains a clinical challenge due to the limited specificity and structure of existing Artificial Intelligence tools. The purpose of this international, multicenter clinical trial is to developed and validated an agent-based framework that integrates vision-language models and large language models to enhance the diagnostic workflow of anterior segment diseases.
Study Design
- Study Type
- Observational
- Observational Model
- Other
- Time Perspective
- Cross Sectional
Eligibility Criteria
- Ages
- 18 Years to — (Adult, Older Adult)
- Sex
- All
- Accepts Healthy Volunteers
- Yes
Inclusion Criteria
- •Informed consent obtained;
- •Participants should be sufficiently able to read, write, and understand Chinese or English;
- •For normal participants: individuals should have no concerns related to their eyes.
- •For participants with eye-related chief complaints: individuals should have specific concerns or issues related to their eyes.
Exclusion Criteria
- •Incomplete clinical data to support final diagnosis;
- •Patients who, in the opinion of the attending physician or clinical study staff, are too medically unstable to participate in the study safely.
Arms & Interventions
Normal participants
Healthy individuals who have no concerns related to their eyes.
Intervention: Multimodal Vision-language Model Diagnosis (Diagnostic Test)
Patients with Eye-related Chief Complaints
Individuals who have specific concerns or issues related to their eyes, which they consider as the main reason for seeking medical attention or making a complaint.
Intervention: Multimodal Vision-language Model Diagnosis (Diagnostic Test)
Outcomes
Primary Outcomes
Diagnostic accuracy of multimodal vision-language model.
Time Frame: from July 2025 to September 2025
For each patient, the diagnoses generated by the multimodal vision-language model and the clinical diagnosis provided by skilled clinicians were documented and compared. Consistency between the two diagnoses indicates the program's precision in clinical practice.
Secondary Outcomes
No secondary outcomes reported
