[論文レビュー] Natural language processing to identify lupus nephritis phenotype in electronic health records
本研究では、電子歴史記録(EHR)内のラウスス・ネフリトス病型を同定する自然言語処理(NLP)ベースのアルゴリズムを開発および検証した。構造化データと病理報告書およびノートからの臨床テキストを統合した。最良の性能を示したモデル—CUI、正規表現、および特徴数の混合を使用—は、 Vanderbilt University Medical Center での外部検証で F-スコア 0.96 を達成し、ベースラインのルールベース手法(F-スコア 0.62)を著しく上回った。
Systemic lupus erythematosus (SLE) is a rare autoimmune disorder characterized by an unpredictable course of flares and remission with diverse manifestations. Lupus nephritis, one of the major disease manifestations of SLE for organ damage and mortality, is a key component of lupus classification criteria. Accurately identifying lupus nephritis in electronic health records (EHRs) would therefore benefit large cohort observational studies and clinical trials where characterization of the patient population is critical for recruitment, study design, and analysis. Lupus nephritis can be recognized through procedure codes and structured data, such as laboratory tests. However, other critical information documenting lupus nephritis, such as histologic reports from kidney biopsies and prior medical history narratives, require sophisticated text processing to mine information from pathology reports and clinical notes. In this study, we developed algorithms to identify lupus nephritis with and without natural language processing (NLP) using EHR data. We developed four algorithms: a rule-based algorithm using only structured data (baseline algorithm) and three algorithms using different NLP models. The three NLP models are based on regularized logistic regression and use different sets of features including positive mention of concept unique identifiers (CUIs), number of appearances of CUIs, and a mixture of three components respectively. The baseline algorithm and the best performed NLP algorithm were external validated on a dataset from Vanderbilt University Medical Center (VUMC). Our best performing NLP model incorporating features from both structured data, regular expression concepts, and mapped CUIs improved F measure in both the NMEDW (0.41 vs 0.79) and VUMC (0.62 vs 0.96) datasets compared to the baseline lupus nephritis algorithm.
研究の動機と目的
- 自然言語処理(NLP)を用いて、EHR におけるラウスス・ネフリトス病型の同定を改善すること。
- EHR における構造化データおよび手術コードにのみ依存する限界を克服すること。
- 特に生検報告書および記述的ノートを含む非構造化臨床テキストを、ラウスス・ネフリトス病型分類に統合すること。
- NLP強化アルゴリズムの性能を Vanderbilt University Medical Center からの外部データセットで検証すること。
- 大規模な観察研究および臨床試験を支援するため、EHR におけるラウスス・ネフリトス病の正確かつスケーラブルな型別分類を可能にすること。
提案手法
- 構造化 EHR データ(例:手術コード、検査結果など)のみを用いたベースラインのルールベースアルゴリズムを開発した。
- 異なる特徴セットを用いた正則化ロジスティック回帰を用いた3つの NLP モデルを構築した:(1) CUI の陽性記載、(2) CUI の頻度、(3) CUI 記載、頻度、および正規表現マッチのハイブリッド。
- 記述的テキスト内の臨床概念を統合医療用語語彙(UMLS)の概念固有識別子(CUI)にマッピングして標準化した。
- 臨床ノートおよび病理報告書内のラウスス・ネフリトス病関連用語を検出するために正規表現を適用した。
- ラベル付き EHR データを用いてロジスティック回帰モデルを訓練およびチューニングし、ラウスス・ネフリトス病を有するか否かを分類するようにした。
- 一般化性を評価するために、独立したデータセット(Vanderbilt University Medical Center:VUMC)を用いて、最も性能の良い NLP モデルを検証した。
実験結果
リサーチクエスチョン
- RQ1構造化データにのみ依存するルールベース手法と比較して、NLP テクニックは EHR におけるラウスス・ネフリトス病型分類の正確性を向上させることができるか?
- RQ2CUI 記載、CUI 頻度、またはハイブリッド特徴のどの組み合わせが、ラウスス・ネフリトス病の検出において最高のパフォーマンスを達成するか?
- RQ3最も性能の良い NLP モデルの性能は、Vanderbilt University Medical Center からの外部で独立した EHR データセットにどの程度一般化されるか?
- RQ4非構造化臨床ノートおよび病理報告書は、構造化データを上回る正確なラウスス・ネフリトス病同定にどの程度寄与するか?
- RQ5言語的パターンと標準化された医療概念(CUI)の両方を組み込んだハイブリッド NLP モデルは、一方の特徴のみを用いたモデルを上回る性能を示せるか?
主な発見
- CUI 記載、CUI 頻度、および正規表現マッチを組み合わせた最良の NLP モデルは、Vanderbilt University Medical Center(VUMC)の外部検証データセットで F-スコア 0.96 を達成した。
- この NLP モデルは、同じ VUMC データセットで F-スコア 0.62 を達成したベースラインのルールベースアルゴリズムを著しく上回った。
- NMEDW データセットでは、最良の NLP モデルが F-スコア 0.79 を達成したのに対し、ベースラインアルゴリズムは 0.41 であった。
- 構造化データと、臨床ノートおよび病理報告書から抽出された NLP 特徴の両方を組み合わせることで、型別分類のパフォーマンスが顕著に向上した。
- CUI および正規表現の使用により、非構造化テキストからのラウスス・ネフリトス病関連概念の効果的抽出が可能になった。
- 結果から、NLP が、診断情報が記述的 EHR コンテンツに埋め込まれている場合の型別分類のギャップを効果的に埋めることが示された。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。