Skip to main content
QUICK REVIEW

[论文解读] Model-assisted cohort selection with bias analysis for generating large-scale cohorts from the EHR for oncology research

Benjamin Birnbaum, Nathan C. Nussbaum|arXiv (Cornell University)|Jan 13, 2020
Machine Learning in Healthcare参考文献 11被引用 200
一句话总结

本论文提出 Model-Assisted Cohort Selection (MACS) 及 Bias Analysis,用以高效生成基于电子病历(EHRs)的肿瘤学队列,展示了高预测性能并且在后续分析中未检测到偏差。

ABSTRACT

Objective Electronic health records (EHRs) are a promising source of data for health outcomes research in oncology. A challenge in using EHR data is that selecting cohorts of patients often requires information in unstructured parts of the record. Machine learning has been used to address this, but even high-performing algorithms may select patients in a non-random manner and bias the resulting cohort. To improve the efficiency of cohort selection while measuring potential bias, we introduce a technique called Model-Assisted Cohort Selection (MACS) with Bias Analysis and apply it to the selection of metastatic breast cancer (mBC) patients. Materials and Methods We trained a model on 17,263 patients using term-frequency inverse-document-frequency (TF-IDF) and logistic regression. We used a test set of 17,292 patients to measure algorithm performance and perform Bias Analysis. We compared the cohort generated by MACS to the cohort that would have been generated without MACS as reference standard, first by comparing distributions of an extensive set of clinical and demographic variables and then by comparing the results of two analyses addressing existing example research questions. Results Our algorithm had an area under the curve (AUC) of 0.976, a sensitivity of 96.0%, and an abstraction efficiency gain of 77.9%. During Bias Analysis, we found no large differences in baseline characteristics and no differences in the example analyses. Conclusion MACS with bias analysis can significantly improve the efficiency of cohort selection on EHR data while instilling confidence that outcomes research performed on the resulting cohort will not be biased.

研究动机与目标

  • 动机使用 EHR 数据进行肿瘤学结局研究,并解决由于非结构化数据导致的非随机队列选择。
  • 开发一个可扩展的队列选择方法,包含偏差评估以增强对后续分析的信任。
  • 将该方法应用于转移性乳腺癌(mBC),以证明效率和偏差控制。

提出的方法

  • 使用 TF-IDF + 逻辑回归模型对 17,263 例患者进行训练,以识别目标队列。
  • 在保留的测试集 17,292 例患者上评估性能。
  • 进行偏差分析,比较 MACS 生成的队列与参考标准在众多临床和人口统计变量上的差异。
  • 比较 MACS 与非 MACS 队列之间变量的分布。
  • 显示 (i) MACS 实现了高辨别度,(ii) 偏差并未显著改变示例分析。

实验结果

研究问题

  • RQ1MACS 是否能提高从 EHR 数据中进行肿瘤学研究的队列选择效率?
  • RQ2相较于参考标准,MACS 是否引入可检测的基线特征偏差?
  • RQ3在 MACS 生成的队列上进行的分析,结果是否与无偏差的参考分析一致?

主要发现

  • AUC 为 0.976,表明 MACS 选择器具有很强的辨别能力。
  • 96.0% 的敏感性,表明对目标队列的高真实阳性捕获率。
  • 抽象效率提升 77.9%,显示显著的工作流程改进。
  • 偏差分析显示 MACS 与参考队列之间的基线特征无显著差异。
  • 在 MACS-derived 队列与参考分析之间未发现差异。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。