Skip to main content
QUICK REVIEW

[论文解读] Towards unraveling calibration biases in medical image analysis

María Agustina Ricci Lara, Candelaria Mosquera|arXiv (Cornell University)|May 9, 2023
Cutaneous Melanoma Detection and Management被引用 5
一句话总结

本文研究了人工智能模型在皮肤病变分类中校准偏差的问题,揭示了常见校准度量指标(如期望校准误差,ECE)会因不同人口统计子群体之间的样本量不平衡(尤其是浅肤色与深肤色患者之间)而系统性地产生偏差。研究表明,此类偏差可能导致虚假的公平性评估,建议在低资源子群体中使用鲁棒指标(如Delta-Brier或Delta-CE)以确保医疗人工智能中可靠公平性审计。

ABSTRACT

In recent years the development of artificial intelligence (AI) systems for automated medical image analysis has gained enormous momentum. At the same time, a large body of work has shown that AI systems can systematically and unfairly discriminate against certain populations in various application scenarios. These two facts have motivated the emergence of algorithmic fairness studies in this field. Most research on healthcare algorithmic fairness to date has focused on the assessment of biases in terms of classical discrimination metrics such as AUC and accuracy. Potential biases in terms of model calibration, however, have only recently begun to be evaluated. This is especially important when working with clinical decision support systems, as predictive uncertainty is key for health professionals to optimally evaluate and combine multiple sources of information. In this work we study discrimination and calibration biases in models trained for automatic detection of malignant dermatological conditions from skin lesions images. Importantly, we show how several typically employed calibration metrics are systematically biased with respect to sample sizes, and how this can lead to erroneous fairness analysis if not taken into consideration. This is of particular relevance to fairness studies, where data imbalance results in drastic sample size differences between demographic sub-groups, which, if not taken into account, can act as confounders.

研究动机与目标

  • 研究不同人口统计子群体之间的样本量不平衡如何影响医学图像分析中校准度量指标的可靠性。
  • 评估在数据不平衡存在的情况下,如期望校准误差(ECE)等常用校准度量指标是否会产生误导性的公平性评估。
  • 识别对有限样本量不敏感的校准度量指标,以实现在代表性不足群体中更准确的公平性审计。
  • 评估模型校准质量与样本量偏差在公平性评估中的相互作用。
  • 倡导使用包容性数据集和鲁棒校准度量指标,以确保皮肤科影像中人工智能的公平性。

提出的方法

  • 在包含不平衡肤色子群体的真实世界皮肤科数据集上训练的皮肤病变分类模型上,评估了多种校准度量指标(如ECE、Brier Score、Delta-Brier、Delta-CE)。
  • 通过控制实验和真实数据,分析了每种度量指标对浅肤色与深肤色患者群体之间样本量差异的敏感性。
  • 使用包含1,494张图像和Fitzpatrick肤色类型标签的PAD-UFES-20数据集,评估不同人口统计子群体间的公平性。
  • 比较了朴素度量应用与重采样及鲁棒度量策略在缓解样本量偏差方面的效果。
  • 研究了模型校准质量与度量指标对样本量敏感性的相互作用,特别是在模型重新校准后的情况。
  • 提出了校准度量指标的两分量解释:一个由有限样本量效应驱动,另一个反映真实的模型校准偏差。
(A)
(A)

实验结果

研究问题

  • RQ1在皮肤科图像分析中,当不同人口统计子群体之间存在极端样本量不平衡时,常见校准度量指标的行为如何?
  • RQ2由于有限样本量效应,像期望校准误差(ECE)这样的度量指标在多大程度上会产生虚假的公平性结论?
  • RQ3哪些校准度量指标受样本量差异影响最小,因而更适合用于低资源子群体的公平性审计?
  • RQ4模型重新校准如何影响校准度量指标对样本量偏差的敏感性?
  • RQ5在数据不平衡的医学影像数据集中,鲁棒度量(如Delta-Brier或Delta-CE)是否能提供更可靠的公平性评估?

主要发现

  • 如期望校准误差(ECE)等常见校准度量指标对样本量高度敏感,并且样本越少,其性能系统性地越差,从而导致不公平性评估偏差。
  • 研究发现,校准度量指标的偏差在已良好校准的模型中最为显著,尤其是在重新校准后,样本量效应占据主导地位。
  • Delta-Brier和Delta-CE等度量指标对样本量的敏感性较低,因此在代表性不足群体的公平性审计中更为可靠。
  • 即使在区分性能(如AUC或准确率)无显著差异的情况下,仍可能因度量指标偏差而非真实模型偏差而出现校准差异。
  • 模型校准质量与样本量之间的相互作用导致度量指标表现出两分量行为:一个由有限样本量效应驱动,另一个由真实的模型校准偏差驱动。
  • 开放数据库中深肤色个体的标注肤色数据稀缺,限制了进行稳健公平性评估的能力,凸显了构建包容性数据集的必要性。
(B)
(B)

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。