Skip to main content
QUICK REVIEW

[論文レビュー] Electronic health record phenotyping improves detection and screening of type 2 diabetes in the general United States population: A cross-sectional, unselected, retrospective study

Ariana Anderson, Wesley T. Kerr|arXiv (Cornell University)|Jan 10, 2015
Diabetes Management and EducationMedicine参考文献 35被引用数 22
ひとこと要約

本研究では、多変量ロジスティック回帰およびランダムフォレストを用いて、診断、薬剤、臨床変数を含む包括的なEHRデータを活用することで、米国一般人口における2型糖尿病の検出が著しく向上することを示している。完全なEHRモデルは、BMI、年齢、性別、ライフスタイル要因のみを用いた従来のリスクモデルよりも優れた性能を示した(p<0.001)、また、片頭痛や心律不全が糖尿病と逆相関することを明らかにした。

ABSTRACT

Objectives: In the United States, 25% of people with type 2 diabetes are undiagnosed. Conventional screening models use limited demographic information to assess risk. We evaluated whether electronic health record (EHR) phenotyping could improve diabetes screening, even when records are incomplete and data are not recorded systematically across patients and practice locations. Methods: In this cross-sectional, retrospective study, data from 9,948 US patients between 2009 and 2012 were used to develop a pre-screening tool to predict current type 2 diabetes, using multivariate logistic regression. We compared (1) a full EHR model containing prescribed medications, diagnoses, and traditional predictive information, (2) a restricted EHR model where medication information was removed, and (3) a conventional model containing only traditional predictive information (BMI, age, gender, hypertensive and smoking status). We additionally used a random-forests classification model to judge whether including additional EHR information could increase the ability to detect patients with Type 2 diabetes on new patient samples. Results: Using a patient's full or restricted EHR to detect diabetes was superior to using basic covariates alone (p&lt;0.001). The random forests model replicated on out-of-bag data. Migraines and cardiac dysrhythmias were negatively associated with type 2 diabetes, while acute bronchitis and herpes zoster were positively associated, among other factors. Conclusions: EHR phenotyping resulted in markedly superior detection of type 2 diabetes in a general US population, could increase the efficiency and accuracy of disease screening, and are capable of picking up signals in real-world records.

研究の動機と目的

  • 米国集団における未診断の2型糖尿病の検出を改善すること。
  • 記録が不完全または一貫性のない場合でも、EHR由来の表現型がスクリーニングの正確性を向上させることを評価すること。
  • EHRベースのモデルと、人口統計学的および基本的臨床要因にのみ依存する従来のリスクモデルの性能を比較すること。
  • 実世界のEHRデータを用いて、2型糖尿病に関連する新たな臨床的関連性を同定すること。
  • ランダムフォレストを用いたアウトオブバッグ予測を用いて、EHR表現型解析の頑健性を検証すること。

提案手法

  • 2009年から2012年までの期間、9,948名の米国患者のEHRデータを用いた横断的・後向き分析を実施した。
  • 診断、薬剤、従来の予測要因を含む完全なEHRデータを用いて、2型糖尿病の現在の状態を予測する多変量ロジスティック回帰モデルを開発した。
  • 薬剤データを除いた制限付きEHRモデルを構築し、その影響が予測性能に与える影響を評価した。
  • BMI、年齢、性別、高血圧、喫煙状態の基本的共変量のみを用いた従来のモデルを構築した。
  • アウトオブバッグデータにおける予測性能を評価し、新たなシグナル関連性を同定するためにランダムフォレスト分類モデルを適用した。
  • 受信者操作特性(ROC)分析を用いて、3つのモデル間の性能を比較した。

実験結果

リサーチクエスチョン

  • RQ1EHR表現型解析は、従来のスクリーニングモデルと比較して、2型糖尿病の検出を著しく改善できるか?
  • RQ2薬剤データの組み込みが、EHRベースの糖尿病スクリーニングモデルの予測正確性に与える影響は何か?
  • RQ3従来のリスク要因を超えた、実世界のEHRデータにおいて2型糖尿病を予測する新たな臨床的関連性は何か?
  • RQ4データが患者や施設間で不完全または一貫性がない場合でも、EHR表現型解析は高い性能を維持できるか?
  • RQ5ランダムフォレストモデルは、EHRベースの表現型解析において、ロジスティック回帰の結果をどの程度再現し一般化できるか?

主な発見

  • 完全なEHRモデルは、従来のモデルよりも2型糖尿病の検出において顕著に優れていた(p < 0.001)、優れた識別性能を示した。
  • 薬剤データを除いた制限付きEHRモデルでさえ、従来のモデルよりも顕著に優れた性能を示した(p < 0.001)、診断および臨床データそのものが検出を向上させることを示した。
  • ランダムフォレストモデルはアウトオブバッグデータにおいて結果をうまく再現し、モデルの頑健性および一般化可能性を確認した。
  • 片頭痛および心律不全は2型糖尿病と負の相関関係にあり、保護的または逆転の臨床的シグナルを示唆した。
  • 急性気管支炎および帯状疱疹は2型糖尿病と正の相関関係にあり、共存疾患または予測的関連性を示唆した。
  • EHR表現型解析は、実世界の非構造的かつ不完全な臨床記録から意味のある生物学的シグナルを効果的に抽出でき、スクリーニングの効率性と正確性を向上させた。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。