Evaluation of Large Language Models for Transforming Structured Dental Radiology Data Into Narrative Radiology Reports
Trial Snapshot
- Phase
- Not Applicable
- Status
- Not yet recruiting
- Sponsor
- Enrollment
- 100
- Locations
- 1
- Primary Endpoint
- Factual consistency of large language model-generated dental radiology reports with structured source data
Study Overview
Brief Summary
The purpose of this observational methodological study is to evaluate whether large language models can transform structured dental radiology data into clear narrative radiology reports. Large language models are computer programs that can generate text from information provided to them. In this study, the input will consist of organized dental radiology findings, such as chart-style or diagram-based information about teeth and surrounding structures.
Dental radiology reports are used by dentists and other health care providers to understand imaging findings and support clinical documentation. Preparing narrative reports may be time-consuming, and the wording of reports may vary between clinicians. This study will examine whether language-model-assisted report generation can produce reports that are complete, accurate, understandable, and clinically useful.
The study will compare reports generated with support from large language models with traditionally prepared reports. Researchers will also assess how the wording of the prompt and selected model parameters influence report quality. In addition, the study will analyze errors and safety risks in generated reports and evaluate whether such a system could be practical in a dental radiology workflow. The language model will not make treatment decisions, and generated reports will be used for research evaluation only.
Detailed Description
This study is designed to evaluate the use of large language models for converting structured dental radiology data into narrative radiology reports. The project focuses on the quality, safety, and practical usability of language-model-assisted report generation in dental radiology.
Structured dental radiology data will be used as the input for the language model. These data may include organized findings recorded in a diagram, chart, or predefined structured format. The model will be asked to transform this structured information into a narrative report resembling a conventional dental radiology description. The study does not evaluate the model as an autonomous diagnostic system. The model will not independently interpret radiographic images, establish a diagnosis, or recommend treatment. Its role is limited to generating narrative text from already structured radiological information.
The study will include several related analyses. First, the investigators will assess whether a large language model can reliably transform structured dental radiology findings into a narrative report. Generated reports will be evaluated for completeness, factual consistency with the source data, clarity, terminology, and clinical readability.
Second, the study will examine how prompt construction and model parameters affect the quality of the generated reports. Different prompt formats and selected generation settings may be compared to identify configurations associated with higher report quality and fewer errors.
Third, reports generated with model assistance will be compared with traditionally prepared narrative reports. The comparison may include blinded assessment by qualified evaluators, who will judge report quality without knowing whether a report was generated traditionally or with model support.
Study Design
- Study Type
- Observational
- Observational Model
- Other
- Time Perspective
- Retrospective
Eligibility Criteria
- Ages
- 18 Years to — (Adult, Older Adult)
- Sex
- All
- Accepts Healthy Volunteers
- No
Inclusion Criteria
- •Dental radiology records based on dental X-ray examination performed on the basis of a written referral from a dentist or physician
- •Dental X-ray examinations performed for screening, diagnostic, or treatment-planning purposes
- •Records from patients with permanent dentition after completion of exfoliation
Exclusion Criteria
- •Records from patients with mixed dentition before completion of exfoliation
- •Records with incomplete, ambiguous, or internally inconsistent structured dental radiology data preventing reliable report generation
- •Records with missing information required for evaluation of the generated report
- •Duplicate records from the same radiographic examination
- •Records in which anonymization or pseudonymization cannot be ensured
Outcomes
Primary Outcomes
Factual consistency of large language model-generated dental radiology reports with structured source data
Time Frame: At the time of report generation and expert evaluation, up to 12 months
Factual consistency will be assessed by comparing each large language model-generated narrative dental radiology report with the corresponding structured dental radiology source data. Expert evaluators will assess whether the generated report accurately reflects the source data without adding findings, omitting findings, changing tooth numbering, or altering the clinical meaning of the structured findings. The outcome will be reported as the proportion of generated reports without clinically relevant factual inconsistency and/or as the number and type of factual inconsistencies per report.
Secondary Outcomes
- Completeness of large language model-generated dental radiology reports(At the time of report generation and expert evaluation, up to 12 months)
- Error rate and error categories in large language model-generated dental radiology reports(At the time of report generation and expert evaluation, up to 12 months)
- Overall quality score of dental radiology reports(At the time of blinded or non-blinded expert evaluation, up to 12 months)
- Difference in expert-rated quality between traditional and large language model-assisted dental radiology reports(At the time of comparative expert evaluation, up to 12 months)
- Effect of prompt design and model parameters on generated report quality(At the time of prompt and parameter comparison, up to 12 months)
- Usability of the large language model-assisted dental radiology reporting workflow(At the time of usability assessment, up to 12 months)
