跳至主要内容
临床试验/NCT07795645
NCT07795645招募中不适用

Development and Validation of Fine-Tuning Techniques for Large Models in Gastric, Cardiovascular and Cerebrovascular Diseases

Beijing Friendship Hospital1 个研究点 分布在 1 个国家目标入组 96,000 人开始时间: 2024年11月1日最近更新:
适应症
干预措施

试验速览

阶段
不适用
状态
招募中
入组人数
96,000
试验地点
1
主要终点
Establishment of the Platform

研究概览

简要总结

The goal of this observational study is to leverage the abundant patient resources and standardized medical records from Beijing Friendship Hospital, Xuanwu Hospital, and Beijing Anzhen Hospital, combined with the existing data and knowledge platform of guidelines, consensus, medical literature, and dialogue data from Beijing Haitian Ruisheng Science Technology Co.,Ltd, with Beijing Zhilan Medical Technology Co., Ltd. conducting the fine-tuning, optimization, and validation of the medical large language model. The model is fine-tuned according to the consultation and diagnostic needs of different departments to improve the quality and efficiency of hospital medical services, enhance intelligence, and elevate the level of medical care. Through deployment to hospitals at all levels, it aims to achieve standardized services and support graded diagnosis and treatment. The overall research includes medical big data construction, medical knowledge graph construction, medical large model training and fine-tuning, and large model application platform development and deployment.

详细描述

This is a multicenter, ambispective cohort study conducted across three hospitals in Beijing (Beijing Friendship Hospital as the lead site, Xuanwu Hospital, and Beijing Anzhen Hospital), utilizing approximately 80000 retrospective historical medical records from patients diagnosed with chronic gastritis, gastric cancer, gastro esophageal reflux, coronary artery disease, or stroke treated between July 2014 and June 2024, along with approximately 16000 prospectively enrolled patients with the same diseases between July 2024 and June 2027.

  1. Smart Medical Big Data Dataset Construction: this study aims to build a smart medical big data dataset centered on the core concept of integrating abundant medical data and knowledge resources to provide solid support for smart medical research and practice. Through close collaboration with hospitals, the research group has accumulated rich multimodal medical data, including electronic medical records, medical images and reports, laboratory results, and genetic information. In parallel, through long-term research and development, the group has accumulated extensive data resources including medical guidelines, expert consensus, clinical databases, medical literature, medical encyclopedias, medical patents, and doctor-patient dialogue data. According to the stage of the project, datasets will be constructed in phases based on specific needs. The early pre-training phase will require medical general knowledge datasets, with data sources including medical guidelines, textbooks, literature, and historical medical records. The later phase of model fine-tuning and application deployment will require disease-specific datasets, with data sources including doctor-patient dialogues, disease-specific medical records, diagnosis and treatment plans, and follow-up records. Through robust data processing capabilities, techniques including data cleaning, standardization, transformation, annotation, and augmentation are employed to convert raw data into forms suitable for different application scenarios. This process not only provides comprehensive training data support for the multi-disease medical large model but also supports future diversified application scenarios ranging from clinical diagnosis support to disease prediction and personalized medicine. The data processing workflow not only requires efficient handling of data heterogeneity but also promotes the establishment of a unified and efficient data representation and processing framework, ensuring long-term data accuracy and usability. Given the high complexity and specialization of medical data, the platform will adopt a human-machine collaboration strategy to reduce data annotation costs, combining the expertise of medical specialists with artificial intelligence capabilities to jointly address data annotation and knowledge extraction challenges. Through continuous input of expert knowledge and ongoing training and optimization, the AI models can improve data processing accuracy and efficiency, thereby providing a solid data foundation for the research, development, and practical application of multi-disease medical large models.
  2. Large-Scale Medical Knowledge Graph Construction: large-scale, high-quality medical knowledge graphs are one of the key data components for achieving fine tuning and application of medical large models. As a core technology, medical knowledge graphs are progressively demonstrating their immense potential and value. A medical knowledge graph is a structured knowledge base that integrates a vast amount of knowledge from the medical field, including diseases, drugs, treatment methods, and medical terminology, and establishes the relationships among them. Such a graph can not only provide physicians with comprehensive and accurate medical knowledge but also assist in medical diagnosis and improve the quality of care.To achieve the fine tuning and application of medical large models, the construction of large scale, high-quality medical knowledge graphs is particularly important. Through in depth research, this study plans to extract multimodal knowledge graphs from data sources including guidelines, consensus, and literature. These data sources possess a high degree of authority and professionalism, providing a solid foundation for knowledge graph construction. In the process of building the knowledge graph, this study has adopted the QLora training framework. QLora is a deep learning based natural language processing framework with robust capabilities in entity recognition, relation extraction, and attribute extraction, enabling the construction of a tens of millions level medical knowledge graph. The knowledge graph construction process involves multiple steps. First, medical knowledge modeling is carried out to define the entities, relations, and attributes to be included in the knowledge graph. Second, through entity recognition models, this study can accurately identify medical entities from textual data, such as disease names and drug names. Third, relation extraction models can analyze the relationships among entities, such as the relationship between diseases and symptoms, and the relationship between drugs and treatment methods. Fourth, attribute extraction models can further extract attribute information of entities, such as the incidence of diseases and the dosage of drugs. Finally, through knowledge fusion and completion, the coverage of knowledge is further enhanced, thereby constructing a comprehensive and accurate medical knowledge graph. This graph not only covers various aspects of the medical field but also establishes relationships and attributes among entities, providing critical data support for the fine tuning and validation of medical large models. It is worth noting that the construction of this knowledge graph also fully considers the integration of multimodal data. In practical applications, medical data often exist in multiple forms, such as text, images, and audio. Therefore, in constructing the knowledge graph, this study not only extracts knowledge from textual data but also incorporates technologies such as image recognition and speech recognition to achieve multimodal data integration. This makes the knowledge graph richer and more diverse, enabling it to more comprehensively reflect the actual conditions of the medical field.
  3. Multi-Disease Medical Large Model Fine-Tuning Platform Construction and Fine-Tuning: this study is dedicated to the in-depth exploration and development of large model fine-tuning techniques applicable to different disease areas in the medical field, as well as a fine-tuning platform that integrates expert experience. The research focuses on customized fine-tuning of vertical domain models, while the fine-tuning platform will incorporate feedback from medical experts to ensure professionalism and generalizability. Through a series of experiments and continuous performance optimization, the study aims to meet the requirements of strong specialization and high accuracy in the medical field, addressing the significant challenges posed by variability across different diseases. In addition, the study will thoroughly evaluate the performance of large language models in natural language processing tasks within the medical domain to fully understand their capabilities in real-world applications. This study uses the ChatGLM large model from Zhipu AI as the foundation model. To effectively reduce the computational complexity and memory requirements of model fine-tuning and inference, techniques including quantization and gradient checkpointing will be explored, maintaining model performance while reducing storage footprint and accelerating inference, and effectively balancing the use of computational resources. In the continued pre-training phase, this study leverages large-scale medical datasets and knowledge graphs to construct training data, and trains through language modeling tasks to enable the model to better understand and process medical information. A series of language modeling tasks tailored to the medical domain are designed. Through iterative optimization, the model will gradually acquire capabilities in medical terminology parsing, symptom understanding, disease trend analysis, and treatment recommendation generation, forming the study's general medical large model. Meanwhile, parameter-efficient fine-tuning techniques are adopted to reduce the number of parameters that need to be updated during fine-tuning, lowering computational resource consumption and shortening training time to better adapt to the specific requirements of different diseases. In the model fine-tuning phase, physicians will provide supervisory signals through professional annotation to help the model learn medical domain knowledge and logical reasoning capabilities. This study first constructs a standardized medical dataset containing patient demographic information, symptoms, signs, laboratory and imaging examinations, diagnostic results, and treatment plans. Second, data augmentation techniques including multi-task learning and adversarial training are explored to enhance the model's adaptability, flexibility, and robustness. Meanwhile, prompt-tuning techniques are adopted to generate specific prompts that guide the model to generate accurate disease-relevant answers, enabling the model to provide more precise and targeted outputs when facing different diseases. In the model deployment phase, quantization strategies are adopted to convert floating-point parameters in the model to low-precision integer representations, reducing model size and improving inference speed. The quantized model substantially reduces GPU memory requirements, enabling large models to run on resource-constrained devices, while inference speed is improved and storage requirements are reduced, lowering deployment and distribution costs. The quality evaluation system is a standardized method for measuring model performance across different medical scenarios, and serves as an essential prerequisite for enhancing model capabilities, algorithm selection, transparency building, trust establishment, and standardized model use. Physicians' expertise and practical experience are critical for guiding model fine-tuning and system development, and collaboration with physicians in developing the evaluation system helps ensure that evaluation criteria align with clinical needs. This study will embed a customized evaluation approach tailored to the medical domain and specific diseases, drawing on existing medical model evaluation frameworks and real-world application scenarios to establish a comprehensive, objective, and effective medical model evaluation system. In addition to standardized medical certification examinations, the study will focus on developing datasets and evaluation frameworks for tasks including initial medical history collection, clinical condition analysis and assessment, and medical record organization, advancing the evaluation of generative language models from theoretical research to practical application.
  4. Medical History Collection System Development: during clinical consultations, patients of different educational levels and age groups vary considerably in their ability to describe their medical conditions, and many patients' emotional states affect the objectivity and accuracy of their symptom descriptions, resulting in low consultation efficiency. This study aims to build an intelligent medical history collection system that achieves intelligent history collection, preliminary etiological analysis, and triage through voice interaction, natural language processing, and large model technologies. In implementation, speech recognition and natural language processing technologies will be used to ensure accurate and efficient collection of patient history information. This includes collecting patients' voice and text input and interacting with patients through multi-round dialogue technology to obtain more detailed history information. Appropriate dialogue strategies will be studied, such as questioning, paraphrasing or interpretation, feeling reflection, and self-disclosure, with attention to patients' emotions to improve patient comfort and satisfaction. In addition, physicians will supervise and validate the model through Reinforcement Learning from Human Feedback (RLHF), which will imitate the bedside manner of excellent clinicians to improve information collection efficiency, reduce dialogue rounds and redundant expressions, and enhance patient experience. After collecting history information, the medical large model will be used to analyze patient information, including extracting key information such as chief complaint, history of present illness, past medical history, personal history, and allergy history. On this basis, RAG combined with guidelines will be used for preliminary triage to recommend appropriate departments for patients, improving consultation efficiency. To validate the effectiveness of the methods proposed in this study, experiments will be conducted to collect medical record data from a large number of real patients and compare it with system-generated medical records. Indicators including history collection completeness and triage accuracy will be evaluated to verify the advantages of the proposed methods in practical applications. This study is expected to make important breakthroughs in intelligent medical history collection, providing an efficient and convenient intelligent consultation assistance system for China's medical industry, improving consultation efficiency and reducing the burden on patients and physicians.
  5. Clinical Auxiliary Diagnosis Research: this study aims to build an intelligent clinical auxiliary diagnosis system, with its main purpose being to improve the efficiency of physician consultations, reduce the pressure of differential diagnosis, and effectively decrease the occurrence of missed diagnoses and misdiagnoses. The system has the following key functions: multi-source data fusion, auxiliary diagnosis suggestion generation, and physician feedback collection and fine-tuning. Multi-source data integration is an important aspect of this study. To this end, the study will investigate how to upload and integrate multi-source heterogeneous medical record, laboratory report, and imaging text report data for unified processing within the large model. This study will utilize Optical Character Recognition (OCR) technology optimized for medical documents, enabling it to process complex mixtures of numbers and Chinese-English text. On top of text information extraction, natural language processing techniques will be employed to ensure accurate interpretation of medical terminology and complete contextual understanding, guaranteeing that data is accurate and complete, thus laying a solid foundation for the next step of information extraction. Through the chain-of-thought (CoT) and reasoning capabilities of large language deep learning models, the system is able to simulate physicians' diagnostic thinking and analyze medical conditions. Specifically, the system obtains key patient information through medical record extraction models, including chief complaint, history of present illness, physical examination findings, specialty information, and auxiliary examination results. The model, fine-tuned on a professional medical knowledge base, is capable of cross-disciplinary etiological reasoning and scenario simulation, providing comprehensive and in-depth condition assessments. On the basis of condition analysis, the system will drive knowledge graph retrieval and derive evidence-based diagnostic recommendations from the retrieval results. The knowledge graph structures large amounts of medical data, medical record analyses, and diagnosis and treatment pathways, thereby supporting the system to consider complex medical relationships and possibilities when generating diagnostic suggestions. For fine-tuning, physicians can provide feedback on the system's auxiliary diagnosis suggestions through a data annotation interface, further improving the system's applicability, accuracy, and reliability. In addition, the system will support customization settings to meet the individual needs of different physicians. In summary, this study is dedicated to building an intelligent auxiliary diagnosis suggestion system to improve the efficiency and accuracy of physician consultations. To achieve this goal, breakthroughs are needed in professional information recognition and extraction, as well as the reasonableness and comprehensiveness of medical record analysis. It is believed that through the research and implementation of this study, positive impacts will be brought to the medical field, improving the quality and level of medical services.
  6. Medical Record Generation and Quality Control Research: this study aims to explore the application of large language models in standardized medical record generation and quality assessment. Medical record quality has a direct impact on medical diagnosis, medical quality management, and medical research, and the requirements and quality standards for medical record writing vary significantly across different hospitals, departments, and physicians. This study will build a medical record generation system that achieves functions including automatic entry of patient information, automatic medical record generation, physician adjustment of records, and medical record standardization assessment, to improve the standardization and quality of medical records. In terms of real-time entry of patient information, the system allows physicians to upload consultation audio, medical records, laboratory and examination reports, and imaging text reports. In terms of automatic medical record generation, this study will apply large language models to complete natural language processing tasks that play key roles in medical record generation, such as entity recognition, text classification, and semantic understanding. This technical approach will enable the large model to more accurately extract key information from physicians' voice or text input, forming medical records that comply with writing standards (including chief complaint, history of present illness, past medical history, personal history, marital and reproductive history, family history, and auxiliary examinations). In terms of medical record quality assessment, this study will define the evaluation criteria for medical record quality, which will be constructed based on medical record industry standards as well as physicians' adjustments and feedback during use. In terms of medical record standardization prompting, this study will conduct in-depth analysis of high-quality medical records and enable the model to learn the characteristics of high-quality records during fine-tuning through contrastive learning, such as detailed history descriptions, accurate symptom documentation, and reasonable treatment plans. On this basis, the study will explore the application of large language models in medical record completion or correction, such as identifying missing key information in records and prompting physicians to supplement it, or detecting potential logical errors and contradictions in records and proposing modification suggestions. The medical record generation system includes an interactive interface that can visually display patient information in different fields, such as chief complaint, history of present illness, and past medical history. Physicians can manually adjust the information before submission, and the modified data will serve as training data for the medical record generation module for continuous model optimization. Meanwhile, the medical record quality control system automatically detects generated records. When core fields are missing, it will remind physicians to make necessary additions and modifications through highlighting, submission blocking, and other methods. Through the application of large language models in medical record generation and quality assessment, this study is expected to improve the standardization and quality of medical records, thereby enhancing the level and efficiency of medical services. In addition, this study will help reduce the burden on physicians in medical record writing, enabling them to focus more on patient treatment and rehabilitation. The results of this study are expected to provide a new solution for the medical industry to address current issues in medical record quality and standardization.
  7. Doctor-Patient Consultation Audio Data Collection and Processing Research This study aims to provide a solution for consultation audio data collection and processing for research areas including intelligent medical history collection, rapid triage, and clinical auxiliary diagnosis. Through efficient collection and precise processing of doctor-patient consultation dialogue audio, the rich consultation audio data will be transformed into structured information, facilitating in-depth understanding of clinical needs and patient conditions, providing training data for large models, and thereby advancing the development of intelligent healthcare and precision medicine. The collection and processing of doctor-patient consultation audio data face multiple challenges, one of which is the highly heterogeneous speech recognition problem. Doctor-patient dialogues are filled with specialized terminology, dialects, different accents, and speech characteristics of elderly patients, all of which significantly increase the difficulty of speech recognition. In addition, noise interference in the consultation environment, non-standard expressions in dialogues, and rapid dialogue exchanges also pose challenges to accurate transcription. To address these challenges, more refined speech processing algorithms need to be studied, deep learning and other algorithmic techniques need to be introduced, and the knowledge of medical domain experts needs to be incorporated to improve recognition accuracy and robustness. Furthermore, protecting patient privacy and data security is another key research consideration. It must be ensured that strict data protection measures are taken throughout the entire process of collection, transmission, storage, and processing of consultation audio data, and that relevant laws and regulations are complied with, using appropriate encryption and anonymization technologies to guarantee the security of patient information. The goal of this study is to efficiently collect audio data from the consultation process and convert this data into accurate textual content. On this basis, the study will also explore preliminary analysis of the collected dialogue content, such as identifying key medical information (symptoms, medications, treatment processes, etc.), to better organize and utilize these data. Although in-depth medical analysis and diagnostic suggestion generation are not within the direct scope of this study, by providing structured and precise dialogue records, this study will lay a solid foundation for these subsequent studies.
  8. Cross Center External Validation: this study aims to validate the generalizability of the intelligent diagnosis and treatment systems for gastric cancer, chronic gastritis, gastro esophageal reflux, coronary artery disease, and stroke. The fine tuned intelligent diagnosis and treatment systems will be transferred from the training center to other centers for external cross validation. By comparing the information collected from large model patient interactions and physician patient interactions, this study will evaluate the feasibility and accuracy of applying the fine tuned intelligent diagnosis and treatment systems in different center environments, thereby validating the generalizability of the fine tuning technology.

研究设计

研究类型
Observational
观察模型
Cohort
时间视角
Other

入排标准

年龄范围
18 Years 至 —(Adult, Older Adult)
性别
All
接受健康志愿者

入选标准

  • Retrospective Historical Medical Records: patients with stomach, cardiovascular and cerebrovascular diseases from September 2014 to August 2024 were enrolled.
  • (1) Age ≥ 18 years; (2) Diagnosed with any of the following: chronic gastritis, gastric cancer, gastro esophageal reflux, coronary artery disease, or stroke.
  • Prospective Historical Medical Records: Patients with stomach, cardiovascular and cerebrovascular diseases from September 2024 and August 2027 are enrolled.Outpatient and emergency records are used for large model training, and inpatient records are used for both training and internal validation. Inpatient records are allocated to the training set and internal validation set at a 3:1 ratio. Block randomization is used to reduce bias. Each sample is assigned a raw random number uniformly distributed between 0 and
  • Under the block design, each block contains 4 samples. Within each block, samples are ranked by the raw random number and assigned a random code from 1 to
  • The randomization schedule is prepared by a statistician on a computer system before the start of the study and printed on opaque, sealed envelopes.
  • Age ≥ 18 years; (2) Clinically diagnosed with one of the following: chronic gastritis, gastric cancer, gastro esophageal reflux, coronary artery disease, or stroke; (3) Patient or legally authorized representative able to understand the study and provide informed consent; (4) Clinically stable and able to complete the study procedures

排除标准

  • Retrospective Historical Medical Records:
  • (1) Records with information that cannot be correctly read due to modification or smudging; (2) Examination reports that are smudged or damaged, making them uninterpretable by the large model.
  • Prospective Historical Medical Records:
  • (1) Severe psychiatric disorders (e.g., depression, mania, epilepsy, schizophrenia); (2) Judged by the investigator to be unable to comply with study procedures; (3) Poor audio quality due to accent or recording issues that prevents accurate data capture; (4) Laboratory or imaging reports that are smudged or damaged, making them uninterpretable by the model
  • Withdrawal Criteria:
  • Prospective Historical Medical Records:
  • Participant requests to withdraw from the study during the research process
  • Investigator determines that the study should be terminated based on consideration of the participant's best interests

研究组 & 干预措施

Stomach Diseases

Patients clinically diagnosed with chronic gastritis, gastric cancer or gastro esophageal reflux.

干预措施: Smart Medical Big Data Dataset Construction (Other)

Stomach Diseases

Patients clinically diagnosed with chronic gastritis, gastric cancer or gastro esophageal reflux.

干预措施: Large-Scale Medical Knowledge Graph Construction (Other)

Stomach Diseases

Patients clinically diagnosed with chronic gastritis, gastric cancer or gastro esophageal reflux.

干预措施: Multi-Disease Medical Large Model Fine-Tuning Platform Construction and Fine-Tuning (Other)

Stomach Diseases

Patients clinically diagnosed with chronic gastritis, gastric cancer or gastro esophageal reflux.

干预措施: Medical History Collection System Development (Other)

Stomach Diseases

Patients clinically diagnosed with chronic gastritis, gastric cancer or gastro esophageal reflux.

干预措施: Clinical Auxiliary Diagnosis Research (Other)

Stomach Diseases

Patients clinically diagnosed with chronic gastritis, gastric cancer or gastro esophageal reflux.

干预措施: Medical Record Generation and Quality Control Research (Other)

Stomach Diseases

Patients clinically diagnosed with chronic gastritis, gastric cancer or gastro esophageal reflux.

干预措施: Doctor-Patient Consultation Audio Data Collection and Processing Research (Other)

Stomach Diseases

Patients clinically diagnosed with chronic gastritis, gastric cancer or gastro esophageal reflux.

干预措施: Cross-Center External Validation (Other)

Cardiovascular and Cerebrovascular Diseases

Patients clinically diagnosed with coronary artery disease or stroke.

干预措施: Smart Medical Big Data Dataset Construction (Other)

Cardiovascular and Cerebrovascular Diseases

Patients clinically diagnosed with coronary artery disease or stroke.

干预措施: Large-Scale Medical Knowledge Graph Construction (Other)

Cardiovascular and Cerebrovascular Diseases

Patients clinically diagnosed with coronary artery disease or stroke.

干预措施: Multi-Disease Medical Large Model Fine-Tuning Platform Construction and Fine-Tuning (Other)

Cardiovascular and Cerebrovascular Diseases

Patients clinically diagnosed with coronary artery disease or stroke.

干预措施: Medical History Collection System Development (Other)

Cardiovascular and Cerebrovascular Diseases

Patients clinically diagnosed with coronary artery disease or stroke.

干预措施: Clinical Auxiliary Diagnosis Research (Other)

Cardiovascular and Cerebrovascular Diseases

Patients clinically diagnosed with coronary artery disease or stroke.

干预措施: Medical Record Generation and Quality Control Research (Other)

Cardiovascular and Cerebrovascular Diseases

Patients clinically diagnosed with coronary artery disease or stroke.

干预措施: Doctor-Patient Consultation Audio Data Collection and Processing Research (Other)

Cardiovascular and Cerebrovascular Diseases

Patients clinically diagnosed with coronary artery disease or stroke.

干预措施: Cross-Center External Validation (Other)

结局指标

主要结局

Establishment of the Platform

时间窗: From enrollment to completion of the platform development, assessed up to 2 years

The medical large language model fine-tuning technology platform is established based on the ChatGLM foundation model. The overall research includes medical big data construction, medical knowledge graph construction, medical large model training and fine-tuning, and large model application platform development and deployment.

次要结局

  • Establishment of the Medical Dialogue Audio Collection System(From enrollment to completion of the platform development, assessed up to 2 years)
  • Establishment of the Medical Big Data Platform Database(From enrollment to completion of the platform development, assessed up to 2 years)
  • Establishment of the Large Model Fine-Tuning Technology(From enrollment to completion of the platform development, assessed up to 2 years)
  • Establishment of the Disease-Specific Knowledge Graph Construction Technology(From enrollment to completion of the knowledge graph construction, assessed up to 2 years)
  • Establishment of the Medical History Collection System(From enrollment to completion of the platform development, assessed up to 2 years)
  • Establishment of the Diagnostic Support System(From enrollment to completion of the platform development, assessed up to 2 years)
  • Establishment of the Medical Record Generation and Quality Control System(From enrollment to completion of the platform development, assessed up to 2 years)
  • Establishment of the Cross-Center External Validation(From enrollment to completion of the cross-center clinical validation, assessed up to 3 years)
  • Innovation in Disease-Specific Diagnosis and Treatment Services Based on the Large Model(From enrollment to completion of the platform development, assessed up to 2 years)

研究者

申办方类型
Other
责任方
Sponsor

研究点 (1)

Loading locations...

相似试验