[论文解读] The Medical Deconfounder: Assessing Treatment Effects with Electronic Health Records
医疗去混杂因子法是一种机器学习方法,通过从多药物处方模式中构建替代混杂因子,以调整未观测到的混杂因素,从而从电子健康记录(EHRs)中估计治疗效果。该方法提高了因果推断的准确性,识别出的药物治疗效果与医学文献更为一致,优于经典方法。
The treatment effects of medications play a key role in guiding medical prescriptions. They are usually assessed with randomized controlled trials (RCTs), which are expensive. Recently, large-scale electronic health records (EHRs) have become available, opening up new opportunities for more cost-effective assessments. However, assessing a treatment effect from EHRs is challenging: it is biased by unobserved confounders, unmeasured variables that affect both patients' medical prescription and their outcome, e.g. the patients' social economic status. To adjust for unobserved confounders, we develop the medical deconfounder, a machine learning algorithm that unbiasedly estimates treatment effects from EHRs. The medical deconfounder first constructs a substitute confounder by modeling which medications were prescribed to each patient; this substitute confounder is guaranteed to capture all multi-medication confounders, observed or unobserved (arXiv:1805.06826). It then uses this substitute confounder to adjust for the confounding bias in the analysis. We validate the medical deconfounder on two simulated and two real medical data sets. Compared to classical approaches, the medical deconfounder produces closer-to-truth treatment effect estimates; it also identifies effective medications that are more consistent with the findings in the medical literature.
研究动机与目标
- 解决由于未观测混杂因素(如社会经济状况)导致的从EHR中估计治疗效果时产生的偏差。
- 开发一种可扩展的机器学习方法,无需显式测量混杂因素即可调整多药物混杂因素。
- 与经典因果推断方法相比,提高真实世界观察性EHR数据中治疗效果估计的准确性。
- 识别出对临床结局(如糖尿病患者的HbA1c水平)具有因果效应的药物,其结果与既定医学文献一致。
提出的方法
- 该方法使用概率因子模型对患者的药物记录进行建模,以推断一个替代混杂因子,该因子捕捉所有多药物混杂因素(无论是否观测到)。
- 替代混杂因子从联合药物处方模式中构建,利用了药物使用相关性反映潜在未测量混杂因素的事实。
- 随后拟合一个多元结果模型以估计治疗效果,同时对所开药物和替代混杂因子进行条件处理,以调整混杂偏差。
- 该方法基于先前研究(Wang和Blei,2018)建立的理论保证:替代混杂因子能捕捉所有由多药物使用引起的混杂效应。
- 该方法使用模拟数据集和来自2型糖尿病患者队列的真实EHR数据进行了验证,结果测量指标为HbA1c水平。
实验结果
研究问题
- RQ1当未观测混杂因素(如社会经济状况)导致标准比较产生偏差时,机器学习方法是否能有效估计EHR中的治疗效果?
- RQ2建模联合药物处方模式是否能产生一个比传统调整方法更有效地捕捉未观测混杂因素的替代混杂因子?
- RQ3在真实世界的EHR数据中,医疗去混杂因子法的治疗效果估计与未调整模型和经典因果推断方法相比如何?
- RQ4医疗去混杂因子法识别出的因果药物是否与医学文献中的发现一致?
主要发现
- 在两个模拟数据集中,医疗去混杂因子法产生的治疗效果估计值比经典方法更接近真实值,表明其对未观测混杂因素的偏差有更小的影响。
- 在2型糖尿病患者的真实EHR数据集中,医疗去混杂因子法识别出三种具有因果效应的药物——他克莫司、氨氯地平和氢氯噻嗪——其对HbA1c具有正向影响,这与这些药物对葡萄糖代谢的已知不良影响一致。
- 两种药物——对乙酰氨基酚和阿托伐他汀——被未调整模型错误地识别为具有因果效应,但医疗去混杂因子法正确地将其判定为非因果,表明其特异性得到提升。
- 在去混杂后,医疗去混杂因子法对阿托伐他汀的治疗效果估计值更偏正向,与该药物可能升高血糖的已知特性一致,表明其效果估计偏差更小。
- 该方法成功识别出胰岛素和二甲双胍为具有负向影响HbA1c的因果药物,这与它们在糖尿病中降低血糖的公认作用一致。
- 该方法识别出异丙醇和羟考酮可能与HbA1c相关,提示可能存在新颖发现,需进一步研究。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。