Skip to main content
QUICK REVIEW

[论文解读] Using massive health insurance claims data to predict very high-cost claimants: a machine learning approach

José M. Maisog, Wenhong Li|arXiv (Cornell University)|Dec 30, 2019
Artificial Intelligence in Healthcare参考文献 42被引用 9
一句话总结

本研究利用4800万患者的索赔数据,开发了一项机器学习模型,用于预测极高的医疗费用索赔者(HiCCs),即年费用超过25万美元的患者。最佳模型在0.99风险阈值下AUC达到91.2%,精确率74%,可实现针对高风险患者的精准护理管理,预计每年可节省730万美元净成本。

ABSTRACT

Due to escalating healthcare costs, accurately predicting which patients will incur high costs is an important task for payers and providers of healthcare. High-cost claimants (HiCCs) are patients who have annual costs above $\$250,000$ and who represent just 0.16% of the insured population but currently account for 9% of all healthcare costs. In this study, we aimed to develop a high-performance algorithm to predict HiCCs to inform a novel care management system. Using health insurance claims from 48 million people and augmented with census data, we applied machine learning to train binary classification models to calculate the personal risk of HiCC. To train the models, we developed a platform starting with 6,006 variables across all clinical and demographic dimensions and constructed over one hundred candidate models. The best model achieved an area under the receiver operating characteristic curve of 91.2%. The model exceeds the highest published performance (84%) and remains high for patients with no prior history of high-cost status (89%), who have less than a full year of enrollment (87%), or lack pharmacy claims data (88%). It attains an area under the precision-recall curve of 23.1%, and precision of 74% at a threshold of 0.99. A care management program enrolling 500 people with the highest HiCC risk is expected to treat 199 true HiCCs and generate a net savings of $\$7.3$ million per year. Our results demonstrate that high-performing predictive models can be constructed using claims data and publicly available data alone, even for rare high-cost claimants exceeding $\$250,000$. Our model demonstrates the transformational power of machine learning and artificial intelligence in care management, which would allow healthcare payers and providers to introduce the next generation of care management programs.

研究动机与目标

  • 开发一个高性能预测模型,用于在大规模参保人群中识别极高的医疗费用索赔者(HiCCs)。
  • 通过利用大规模索赔和人口普查数据,改进现有方法,以提高对罕见高成本患者的预测准确性。
  • 通过识别年医疗费用超过25万美元的患者中风险最高的群体,实现主动护理管理。
  • 评估模型在具有挑战性的子人群中的表现,包括参保时间较短、无处方药数据或无既往高成本历史的患者。
  • 估算基于风险的护理管理项目部署后的潜在财务影响。

提出的方法

  • 基于4800万患者的临床、人口统计和索赔特征共6006个变量,构建预测模型。
  • 通过公开的人口普查数据扩充索赔数据,以丰富患者层面的特征。
  • 使用监督机器学习技术训练并比较100多个二分类模型。
  • 通过AUC-ROC和精确率-召回率曲线指标优化模型性能,重点关注高特异性和高精确率。
  • 基于AUC-ROC(91.2%)和在高风险阈值0.99下的精确率,选择表现最佳的模型。
  • 通过将前500名高风险患者纳入模拟护理管理项目,估算净财务节省。

实验结果

研究问题

  • RQ1基于大规模索赔和人口普查数据训练的机器学习模型,能否准确预测年医疗费用超过25万美元的患者?
  • RQ2在数据有限的子人群(如参保时间不足一年或无处方药索赔的患者)中,模型性能如何变化?
  • RQ3基于高风险预测部署目标护理管理项目的潜在财务影响是什么?
  • RQ4与以往发表的HiCC预测方法相比,该模型的性能如何?
  • RQ5在高风险阈值(如0.99)下,该模型能否保持高精确率,以最小化临床干预中的假阳性?

主要发现

  • 最佳机器学习模型的受试者工作特征曲线下面积(AUC)达到91.2%,显著优于此前报道的最高基准值84%。
  • 模型在具有挑战性的子群体中仍保持强劲表现:无既往高成本状态患者的AUC为89%,参保时间不足一年的患者为87%,无处方药数据的患者为88%。
  • 在0.99风险阈值下,模型精确率达到74%,表明74%被标记为高风险的患者确实是真正的HiCCs。
  • 模型在精确率-召回率曲线下面积达到23.1%,反映出在识别罕见事件时具备高精确率的优异表现。
  • 模拟护理管理项目中,若纳入前500名高风险患者,可识别出199名真实HiCCs,并预计每年实现730万美元的净节省。
  • 结果表明,仅使用索赔数据和公开可获取的数据,即可构建高性能的罕见高成本患者预测模型,无需临床记录或电子健康记录(EHR)数据。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。