[论文解读] The Impact of Automatic Pre-annotation in Clinical Note Data Element Extraction - the CLEAN Tool
本文介绍了 CLEAN,一种基于预标注的临床笔记标注系统,可提高从临床笔记中提取数据元素的准确性。通过集成管道(CLEAN-EP)和自定义标注工具(CLEAN-AT),CLEAN 在 F1 分数上显著优于 BRAT(0.896 vs. 0.820),且标注时间无显著差异,证明其在不牺牲效率的前提下提升了准确性与用户满意度。
Objective. Annotation is expensive but essential for clinical note review and clinical natural language processing (cNLP). However, the extent to which computer-generated pre-annotation is beneficial to human annotation is still an open question. Our study introduces CLEAN (CLinical note rEview and ANnotation), a pre-annotation-based cNLP annotation system to improve clinical note annotation of data elements, and comprehensively compares CLEAN with the widely-used annotation system Brat Rapid Annotation Tool (BRAT). Materials and Methods. CLEAN includes an ensemble pipeline (CLEAN-EP) with a newly developed annotation tool (CLEAN-AT). A domain expert and a novice user/annotator participated in a comparative usability test by tagging 87 data elements related to Congestive Heart Failure (CHF) and Kawasaki Disease (KD) cohorts in 84 public notes. Results. CLEAN achieved higher note-level F1-score (0.896) over BRAT (0.820), with significant difference in correctness (P-value < 0.001), and the mostly related factor being system/software (P-value < 0.001). No significant difference (P-value 0.188) in annotation time was observed between CLEAN (7.262 minutes/note) and BRAT (8.286 minutes/note). The difference was mostly associated with note length (P-value < 0.001) and system/software (P-value 0.013). The expert reported CLEAN to be useful/satisfactory, while the novice reported slight improvements. Discussion. CLEAN improves the correctness of annotation and increases usefulness/satisfaction with the same level of efficiency. Limitations include untested impact of pre-annotation correctness rate, small sample size, small user size, and restrictedly validated gold standard. Conclusion. CLEAN with pre-annotation can be beneficial for an expert to deal with complex annotation tasks involving numerous and diverse target data elements.
研究动机与目标
- 评估自动预标注对临床笔记数据元素提取准确性与效率的影响。
- 开发并测试一种新型标注系统 CLEAN,整合预标注以支持人工标注者。
- 将 CLEAN 的性能与可用性与广泛使用的 BRAT 标注工具进行比较。
- 评估预标注是否在不增加时间负担的前提下提升标注正确性与用户满意度。
- 研究系统设计与用户专业程度在标注结果中的作用。
提出的方法
- CLEAN 采用集成管道(CLEAN-EP),结合多个 NLP 模型,为临床笔记生成预标注。
- 自定义标注工具(CLEAN-AT)将预标注呈现给人工标注者,使其能够高效地审查与修正。
- 系统使用来自心力衰竭(CHF)和川崎病(KD)队列的 84 份公开临床笔记进行评估,需提取 87 个数据元素。
- 通过一名领域专家和一名新手标注者,对 CLEAN 与 BRAT 进行了对比可用性测试。
- 性能通过笔记级别 F1 分数、标注时间及用户报告的满意度进行衡量。
- 统计分析评估了系统类型、笔记长度与用户专业程度对正确性与时间的影响。
实验结果
研究问题
- RQ1自动预标注是否提升了临床笔记数据元素提取中人工标注的正确性?
- RQ2CLEAN 与 BRAT 在标注时间与效率方面有何差异?
- RQ3预标注在多大程度上提升了用户满意度与感知有用性?
- RQ4预标注带来的性能提升是否依赖于用户专业程度或笔记复杂度?
- RQ5系统设计、笔记长度或用户类型中,哪个因素对标注正确性与时间的影响最为显著?
主要发现
- CLEAN 的笔记级别 F1 分数为 0.896,显著高于 BRAT 的 0.820(p < 0.001),表明标注正确性得到提升。
- 正确性差异主要归因于所用系统/软件(p < 0.001),而非用户专业程度。
- CLEAN 平均标注时间为每篇 7.262 分钟,BRAT 为 8.286 分钟,差异无显著性(p = 0.188)。
- 笔记长度是影响标注时间的最重要因素(p < 0.001),系统类型也具有显著影响(p = 0.013)。
- 领域专家认为 CLEAN 有用且令人满意,新手标注者则报告其可用性略有提升。
- 结果表明,预标注在不增加时间成本的前提下提升了准确性和可用性,尤其适用于复杂的标注任务。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。