Skip to main content
QUICK REVIEW

[论文解读] Developing a cardiovascular disease risk-factors annotated corpus of Chinese electronic medical records

Jia Su, Bin He|arXiv (Cornell University)|Nov 28, 2016
Biomedical Text Mining and Ontologies参考文献 30被引用 4
一句话总结

本研究基于600份去标识化的中文电子病历(CEMRs),采用临床医生指导的标注指南,从纵向临床数据中提取具有时间属性和断言属性的心血管疾病(CVD)风险因素,构建了一个高可靠性的风险因素标注语料库。该语料库的标注者间一致性F1分数达到0.968,为未来在纵向临床数据中开发CVD风险因素监测与预测系统提供了支持。

ABSTRACT

Objective The goal of this study was to build a corpus of cardiovascular disease (CVD) risk-factor annotations based on Chinese electronic medical records (CEMRs). This corpus is intended to be used to develop a risk-factor information extraction system that, in turn, can be applied as a foundation for the further study of the progress of risk-factors and CVD. Materials and Methods We designed a light-annotation-task to capture CVD-risk-factors with indicators, temporal attributes and assertions explicitly displayed in the records. The task included: 1) preparing data; 2) creating guidelines for capturing annotations (these were created with the help of clinicians); 3) proposing annotation method including building the guidelines draft, training the annotators and updating the guidelines, and corpus construction. Results The outcome of this study was a risk-factor-annotated corpus based on de-identified discharge summaries and progress notes from 600 patients. Built with the help of specialists, this corpus has an inter-annotator agreement (IAA) F1-measure of 0.968, indicating a high reliability. Discussion Our annotations included 12 CVD-risk-factors such as Hypertension and Diabetes. The annotations can be applied as a powerful tool to the management of these chronic diseases and the prediction of CVD. Conclusion Guidelines for capturing CVD-risk-factor annotations from CEMRs were proposed and an annotated corpus was established. The obtained document-level annotations can be applied in future studies to monitor risk-factors and CVD over the long term.

研究动机与目标

  • 为解决中文电子病历(CEMRs)中CVD风险因素提取缺乏标准化标注临床数据的问题。
  • 开发可靠的标注指南,以捕捉具有时间属性和断言状态(如:存在、过去、不存在)的CVD风险因素。
  • 从真实世界的CEMRs中构建大规模、去标识化的CVD风险因素标注语料库,以支持下游NLP应用。
  • 通过实现对风险因素随时间的准确追踪,支持CVD进展的纵向研究。
  • 为开发自动化系统以从临床文本中提取和监控CVD风险因素奠定基础。

提出的方法

  • 设计了一项轻量级标注任务,专注于提取12种预定义的CVD风险因素(如高血压、糖尿病),并明确标注其指示词、时间属性和断言状态。
  • 与临床医生合作,制定了详细的标注指南,包括定义、示例和边界情况。
  • 实施了迭代式标注流程,包括指南起草、标注者培训和指南优化,以确保一致性。
  • 从600份去标识化的出院小结和病程记录中构建了该语料库,确保患者隐私和数据去标识化。
  • 使用F1度量计算标注者间一致性(IAA),以评估标注的可靠性。
  • 利用最终语料库支持临床文本中CVD风险因素信息抽取系统的开发。

实验结果

研究问题

  • RQ1如何在具有时间与断言上下文的前提下,从中文电子病历中可靠地提取CVD风险因素?
  • RQ2在临床NLP任务中,使用临床医生指导的标注指南,可实现多高的标注者间一致性?
  • RQ3能否开发出标准化的标注框架,以支持在多样化临床记录中一致地提取CVD风险因素?
  • RQ4该标注语料库在多大程度上可支持电子健康记录中CVD风险因素的纵向监测?
  • RQ5在中文临床文本中标注CVD风险因素时面临的主要挑战是什么?这些挑战如何通过结构化指南加以缓解?

主要发现

  • 本研究成功从600份去标识化的中文电子病历中构建了风险因素标注语料库,涵盖高血压、糖尿病等12种CVD风险因素。
  • 该语料库实现了高达0.968的标注者间一致性(IAA)F1分数,表明标注具有高度的可靠性与一致性。
  • 临床医生指导的标注指南显著提升了临床文本中风险因素提取的准确性和标准化程度。
  • 时间属性和断言状态(如:存在、过去、不存在)的引入,增强了提取信息的临床实用性。
  • 该标注语料库为训练和评估聚焦于CVD风险因素监测的NLP系统提供了坚实基础。
  • 该语料库适用于CVD进展的长期研究,支持对风险因素随时间变化的追踪。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。