Skip to main content
QUICK REVIEW

[論文レビュー] Using massive health insurance claims data to predict very high-cost claimants: a machine learning approach

José M. Maisog, Wenhong Li|arXiv (Cornell University)|Dec 30, 2019
Artificial Intelligence in Healthcare参考文献 42被引用数 9
ひとこと要約

本研究では、4,800万件の患者の請求データを用いて、年間費用が25万ドルを超えるとされる非常に高コストの請求者(HiCCs)を予測する機械学習モデルを開発した。最高のモデルは、リスク閾値0.99でAUC 91.2%、精度74%を達成し、ターゲットケアマネジメントを可能にし、年間730万ドルの純利益を生み出す可能性がある。

ABSTRACT

Due to escalating healthcare costs, accurately predicting which patients will incur high costs is an important task for payers and providers of healthcare. High-cost claimants (HiCCs) are patients who have annual costs above $\$250,000$ and who represent just 0.16% of the insured population but currently account for 9% of all healthcare costs. In this study, we aimed to develop a high-performance algorithm to predict HiCCs to inform a novel care management system. Using health insurance claims from 48 million people and augmented with census data, we applied machine learning to train binary classification models to calculate the personal risk of HiCC. To train the models, we developed a platform starting with 6,006 variables across all clinical and demographic dimensions and constructed over one hundred candidate models. The best model achieved an area under the receiver operating characteristic curve of 91.2%. The model exceeds the highest published performance (84%) and remains high for patients with no prior history of high-cost status (89%), who have less than a full year of enrollment (87%), or lack pharmacy claims data (88%). It attains an area under the precision-recall curve of 23.1%, and precision of 74% at a threshold of 0.99. A care management program enrolling 500 people with the highest HiCC risk is expected to treat 199 true HiCCs and generate a net savings of $\$7.3$ million per year. Our results demonstrate that high-performing predictive models can be constructed using claims data and publicly available data alone, even for rare high-cost claimants exceeding $\$250,000$. Our model demonstrates the transformational power of machine learning and artificial intelligence in care management, which would allow healthcare payers and providers to introduce the next generation of care management programs.

研究の動機と目的

  • 大規模な被保険者集団において、非常に高コストの請求者(HiCCs)を特定する高精度な予測モデルを開発すること。
  • 大規模な請求データおよびインデックスデータを活用することで、従来の手法を上回り、まれな高コスト患者の予測精度を向上させること。
  • 年間医療費が25万ドルを超える可能性が最も高い患者を事前に同定することで、前向きなケアマネジメントを可能にすること。
  • 短期間の加入履歴、薬剤請求データなし、または過去の高コスト履歴なしのサブグループを含む、挑戦的なサブグループにおけるモデルのパフォーマンスを評価すること。
  • 本モデルを用いたリスクベースのケアマネジメントプログラムを導入した場合の潜在的財務的影響を推定すること。

提案手法

  • 4,800万件の患者の臨床的、人口統計的、請求ベースの特徴を含む6,006の変数を用いて、予測モデルを構築した。
  • 患者レベルの特徴を豊かにするために、公開可能なインデックスデータを請求データに統合した。
  • 教師あり機械学習手法を用いて、100以上の二値分類モデルを訓練・比較した。
  • AUC-ROCおよび精度再現曲線の指標を用いてモデルのパフォーマンスを最適化し、特に特異度と精度に重点を置いた。
  • AUC-ROC(91.2%)と高リスク閾値0.99における精度に基づき、最も優れたパフォーマンスを示したモデルを選定した。
  • 上位500名の高リスク患者を対象にケアマネジメントプログラムをシミュレートし、純財務的利益を推定した。

実験結果

リサーチクエスチョン

  • RQ1大規模な請求データおよびインデックスデータを用いた機械学習モデルは、年間25万ドルを超える医療費を要する患者を正確に予測できるか?
  • RQ2加入期間が短い、または薬剤請求データのない患者など、データが限られたサブグループでは、モデルのパフォーマンスはどのように変化するか?
  • RQ3高リスク予測に基づくターゲットケアマネジメントプログラムを導入した場合の潜在的財務的影響は何か?
  • RQ4本モデルのパフォーマンスは、以前に発表されたHiCC予測手法と比較してどうか?
  • RQ5高リスク閾値(例:0.99)で高い精度を維持できるか。これにより、臨床的介入プログラムにおける誤検出を最小限に抑えることができるか?

主な発見

  • 最高の機械学習モデルは、受信者操作特性曲線下の面積(AUC)が91.2%に達し、以前に報告された最高のベンチマーク(84%)を著しく上回った。
  • 本モデルは、挑戦的なサブグループでも高いパフォーマンスを維持した:過去に高コスト状態でなかった患者ではAUC 89%、1年未満の加入期間の患者では87%、薬剤請求データのない患者では88%であった。
  • リスク閾値0.99で、モデルの精度は74%に達し、高リスクと特定された患者の74%が実際にHiCCであったことを示した。
  • 精度-再現曲線下の面積は23.1%であり、まれなイベントを高い精度で同定できることを示した。
  • 上位500名の高リスク患者を対象にシミュレートしたケアマネジメントプログラムでは、199名の真のHiCCが同定され、年間730万ドルの純利益が見込まれた。
  • 本結果は、臨床ノートや電子健康記録(EHR)を必要とせず、請求データと公開可能なデータのみを用いても、まれな非常に高コスト患者を正確に予測する高パフォーマンスのモデルを構築可能であることを示している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。