[論文レビュー] The Impact of Automatic Pre-annotation in Clinical Note Data Element Extraction - the CLEAN Tool
本論文では、臨床ノートからのデータ要素抽出の正確性を向上させるpre-annotationベースの臨床ノートアノテーションシステムCLEANを紹介する。アンサンブルパイプライン(CLEAN-EP)とカスタムアノテーションツール(CLEAN-AT)を用いたCLEANは、アノテーション時間に顕著な差がなく、F1スコアが著しく向上(0.896 vs. 0.820)した。これは、正確性とユーザ満足度が向上し、効率性を損なわないことを示している。
Objective. Annotation is expensive but essential for clinical note review and clinical natural language processing (cNLP). However, the extent to which computer-generated pre-annotation is beneficial to human annotation is still an open question. Our study introduces CLEAN (CLinical note rEview and ANnotation), a pre-annotation-based cNLP annotation system to improve clinical note annotation of data elements, and comprehensively compares CLEAN with the widely-used annotation system Brat Rapid Annotation Tool (BRAT). Materials and Methods. CLEAN includes an ensemble pipeline (CLEAN-EP) with a newly developed annotation tool (CLEAN-AT). A domain expert and a novice user/annotator participated in a comparative usability test by tagging 87 data elements related to Congestive Heart Failure (CHF) and Kawasaki Disease (KD) cohorts in 84 public notes. Results. CLEAN achieved higher note-level F1-score (0.896) over BRAT (0.820), with significant difference in correctness (P-value < 0.001), and the mostly related factor being system/software (P-value < 0.001). No significant difference (P-value 0.188) in annotation time was observed between CLEAN (7.262 minutes/note) and BRAT (8.286 minutes/note). The difference was mostly associated with note length (P-value < 0.001) and system/software (P-value 0.013). The expert reported CLEAN to be useful/satisfactory, while the novice reported slight improvements. Discussion. CLEAN improves the correctness of annotation and increases usefulness/satisfaction with the same level of efficiency. Limitations include untested impact of pre-annotation correctness rate, small sample size, small user size, and restrictedly validated gold standard. Conclusion. CLEAN with pre-annotation can be beneficial for an expert to deal with complex annotation tasks involving numerous and diverse target data elements.
研究の動機と目的
- 自動pre-annotationが臨床ノートのデータ要素抽出の正確性と効率性に与える影響を評価すること。
- 人間のアノテーターを支援するpre-annotationを統合した新規アノテーションシステムCLEANの開発およびテストを行うこと。
- 広く使用されているBRATアノテーションツールと比較して、CLEANのパフォーマンスと使いやすさを評価すること。
- pre-annotationがアノテーションの正確性とユーザ満足度を向上させるかどうか、時間的負担を増加させないかを評価すること。
- システム設計およびユーザの専門性がアノテーション結果に与える影響を調査すること。
提案手法
- CLEANは、複数のNLPモデルを統合したアンサンブルパイプライン(CLEAN-EP)を用い、臨床ノートのpre-annotationを生成する。
- カスタムアノテーションツール(CLEAN-AT)は、人間のアノテーターにpre-annotationを提示し、効率的に確認・修正できるようにする。
- 評価には、CHFおよびKDコhortの84件の公開臨床ノートを用い、合計87のデータ要素を抽出対象とした。
- 1名のドメインエキスパートと1名の初心者アノテーターを対象に、CLEANとBRATの両方を用いた比較的用の使いやすさテストを実施した。
- パフォーマンスはノート単位のF1スコア、アノテーション時間、およびユーザが報告した満足度で測定した。
- 統計的分析により、システム種別、ノート長、ユーザの専門性が正確性および時間に与える影響を評価した。
実験結果
リサーチクエスチョン
- RQ1自動pre-annotationは、臨床ノートのデータ要素抽出における人間のアノテーションの正確性を向上させるか?
- RQ2CLEANはBRATと比較して、アノテーション時間と効率性にどのような差異を示すか?
- RQ3pre-annotationはどの程度、ユーザ満足度および有用性の認識を向上させるか?
- RQ4pre-annotationによるパフォーマンス向上は、ユーザの専門性やノートの複雑さに依存するか?
- RQ5システム設計、ノート長、ユーザタイプのうち、アノテーション正確性および時間に最も顕著に影響を与える要因は何か?
主な発見
- CLEANはノート単位のF1スコアが0.896を達成し、BRATの0.820よりも顕著に高い(p < 0.001)。これはアノテーション正確性の向上を示している。
- 正確性の差は主に使用したシステム/ソフトウェアに起因する(p < 0.001)、ユーザの専門性によるものではない。
- 平均アノテーション時間は、CLEANが1ノートあたり7.262分、BRATが8.286分であり、有意差は認められなかった(p = 0.188)。
- ノート長がアノテーション時間に最も顕著な要因であった(p < 0.001)、システム種別も影響を与えていた(p = 0.013)。
- エキスパートアノテーターはCLEANを有用で満足できると報告したが、初心者アノテーターは使いやすさにわずかな改善を認めた。
- 結果から、pre-annotationは時間コストを増加させることなく正確性と使いやすさを向上させ、特に複雑なアノテーションタスクにおいて顕著な効果を示すことが示唆された。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。