Skip to main content
QUICK REVIEW

[论文解读] Quantifying Inequality in Underreported Medical Conditions.

Divya Shanmugam, Emma Pierson|arXiv (Cornell University)|Oct 8, 2021
Machine Learning and Data Classification参考文献 31被引用 4
一句话总结

本文提出了一种新颖的方法,通过在协变量偏移假设下使用正样本-未标记样本学习,估算报告不足的医疗状况的相对流行率,从而在无需绝对流行率估计的情况下,实现不同人口群体间的准确比较。该方法在合成数据和真实世界健康数据中均优于基线方法,对协变量偏移假设的轻微违反也表现出鲁棒性。

ABSTRACT

Estimating the prevalence of a medical condition, or the proportion of the population in which it occurs, is a fundamental problem in healthcare and public health. Accurate estimates of the relative prevalence across groups -- capturing, for example, that a condition affects women more frequently than men -- facilitate effective and equitable health policy which prioritizes groups who are disproportionately affected by a condition. However, it is difficult to estimate relative prevalence when a medical condition is underreported. In this work, we provide a method for accurately estimating the relative prevalence of underreported medical conditions, building upon the positive unlabeled learning framework. We show that under the commonly made covariate shift assumption -- i.e., that the probability of having a disease conditional on symptoms remains constant across groups -- we can recover the relative prevalence, even without restrictive assumptions commonly made in positive unlabeled learning and even if it is impossible to recover the absolute prevalence. We provide a suite of experiments on synthetic and real health data that demonstrate our method's ability to recover the relative prevalence more accurately than do baselines, and the method's robustness to plausible violations of the covariate shift assumption.

研究动机与目标

  • 解决在标准报告存在偏差或不完整时,估算报告不足医疗状况在不同人群中的相对流行率的挑战。
  • 开发一种无需估计绝对流行率的方法,因为绝对流行率估计在报告不足的状况中往往不可行。
  • 通过准确识别受某种状况影响更大的群体,实现更公平的健康政策。
  • 在最小假设下运行,特别是避免传统正样本-未标记样本学习中常见的严格条件。
  • 确保在真实世界健康数据中对协变量偏移假设的合理违反具有鲁棒性。

提出的方法

  • 该方法利用正样本-未标记样本学习框架,将确诊病例视为正样本,未确诊病例视为未标记样本。
  • 它假设协变量偏移——即疾病给定症状的条件概率在不同群体间保持不变——从而实现相对流行率的估计。
  • 该方法采用重加权策略,以校正群体特定的报告偏差,从而实现相对流行率比率的估计。
  • 它将估计问题形式化为约束优化任务,以最小化群体间的分布差异。
  • 该方法无需估计绝对流行率,而是专注于群体间相对比例的估计。
  • 通过带校准权重的经验风险最小化,该方法被设计为对协变量偏移假设的适度违反具有鲁棒性。

实验结果

研究问题

  • RQ1我们能否在不假设绝对流行率的情况下,准确估计报告不足医疗状况在不同人口群体中的相对流行率?
  • RQ2与现有正样本-未标记样本学习基线相比,所提出方法在估计相对流行率方面表现如何?
  • RQ3该方法在真实世界健康数据中对协变量偏移假设的违反在多大程度上具有鲁棒性?
  • RQ4当某些人群的诊断系统性地被低估时,该方法能否识别出疾病负担更高的群体?
  • RQ5该方法恢复相对流行率所需的关键假设是什么?与先前工作中的假设相比如何?

主要发现

  • 在合成数据和真实健康数据集上,所提出方法在相对流行率估计方面显著优于基线方法。
  • 即使由于报告不足而无法估计绝对流行率,该方法仍能成功恢复相对流行率比率。
  • 在协变量偏移假设的合理违反下,该方法表现出稳健性能,在现实数据场景中保持了准确性。
  • 与依赖于严格假设(如均匀先验或已知类别不平衡)的现有正样本-未标记样本学习方法相比,该方法表现更优。
  • 实验表明,即使缺乏完整的诊断记录,该方法也能可靠地识别出相对疾病负担更高的群体。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。