Platform for Medical Information Extraction From Incomplete Data
试验速览
- 阶段
- 不适用
- 入组人数
- 10,000
- 试验地点
- 1
- 主要终点
- The number of patients correctly identified by recurrence predictive model
研究概览
简要总结
In order to perform research smoothly, the process of information extraction is required for translating data in clinical text into available format for analysis and statistic. In medical research, the problem of missing data occurs frequently. It is important to develop the method with better imputation performance in the stability and accuracy. The purposes of this project are to provide the data integration and extraction methods for handling the structured and unstructured data sources in more efficient ways, to provide the validation scheme for facilitating the data reviewing of extracted results produced by information extraction modules, to increase the quality of clinical data by comparing the data from different data sources and correcting data errors and inconsistent, to handle the clinical data with the properties of time series and incompleteness, to increase accuracy of data analysis and increase quality of health care by improving the completeness and correctness of clinical data, to provide flexibility of methods in the platform. In the project, the disease topic is focused on the liver cancer patients' clinical data and we hope the methods in the projects can be extended to handle other diseases by replacing these knowledge models in the future.
详细描述
Because of the increasing adoption of Electronic Medical Record (EMR) systems, the data access of EMR is more and more convenient. However, there still have difficulties in analyzing all the clinical data directly due to a large number of records using the narrative format. In order to perform research smoothly, the process of information extraction is required for translating data in clinical text into available format for analysis and statistic. In medical research, the problem of missing data occurs frequently. It is important to develop the method with better imputation performance in the stability and accuracy. The purposes of this project are to provide the data integration and extraction methods for handling the structured and unstructured data sources in more efficient ways, to provide the validation scheme for facilitating the data reviewing of extracted results produced by information extraction modules, to increase the quality of clinical data by comparing the data from different data sources and correcting data errors and inconsistent, to handle the clinical data with the properties of time series and incompleteness, to increase accuracy of data analysis and increase quality of health care by improving the completeness and correctness of clinical data, to provide flexibility of methods in the platform. In the project, the disease topic is focused on the liver cancer patients' clinical data and we hope the methods in the projects can be extended to handle other diseases by replacing these knowledge models in the future.
研究设计
- 研究类型
- Observational
- 时间视角
- Retrospective
入排标准
- 性别
- All
- 接受健康志愿者
- 否
入选标准
- 未提供
排除标准
- 未提供
结局指标
主要结局
The number of patients correctly identified by recurrence predictive model
时间窗: 3 years
The recurrence predictive model is developed using the incomplete data set, this model is used for predicting the recurrent status of patient who received the specific treatment for liver cancer. The number of patients correctly identified by recurrence predictive model is regarded as the primary outcome measure.
次要结局
未报告次要终点
