[论文解读] Selection, Ignorability and Challenges With Causal Fairness
本文挑战了因果公平性模型中独立噪声的假设,指出此类模型——常用于计算反事实公平性——需要类似于随机试验的强 ignorability 假设,而这种假设在现实世界公平性数据中很少成立。作者证明,数据集中的选择偏差会破坏独立噪声假设,从而破坏反事实公平性及一般因果公平性方法。
In this paper we look at popular fairness methods that use causal counterfactuals. These methods capture the intuitive notion that a prediction is fair if it coincides with the prediction that would have been made if someone's race, gender or religion were counterfactually different. In order to achieve this, we must have causal models that are able to capture what someone would be like if we were to counterfactually change these traits. However, we argue that any model that can do this must lie outside the particularly well behaved class that is commonly considered in the fairness literature. This is because in fairness settings, models in this class entail a particularly strong causal assumption, normally only seen in a randomised controlled trial. We argue that in general this is unlikely to hold. Furthermore, we show in many cases it can be explicitly rejected due to the fact that samples are selected from a wider population. We show this creates difficulties for counterfactual fairness as well as for the application of more general causal fairness methods.
研究动机与目标
- 挑战因果公平性中广泛使用的独立噪声模型,这些模型构成了反事实公平性定义的基础。
- 证明这些模型所需的 ignorability 假设在现实世界公平性场景中由于选择偏差而不可信。
- 表明从更广泛总体中进行选择会导致噪声变量之间的依赖,从而违反 i.i.d. 噪声假设。
- 强调反事实公平性与路径特定公平性方法的实际后果。
- 认为当应用于非随机抽样数据时,当前的因果公平性方法在根本上被削弱。
提出的方法
- 使用具有独立噪声的结构因果模型(SCM)来形式化反事实公平性,依赖 do-演算和潜在结果。
- 应用 ignorability 假设以表明在理想条件下,反事实与受保护属性无关。
- 证明选择机制(例如,剔除非工作者)会导致噪声变量与受保护属性之间的依赖。
- 使用图模型和忠实性假设表明,如果独立噪声模型适合完整总体,则无法适合所选样本。
- 使用美国人口普查、德国信贷和法学院数据集的实证数据表明,观察到的性别不平衡与 ignorability 假设相矛盾。
- 应用反事实公平性定义表明,即使一个模型在真实 SCM 下是反事实公平的,由于选择引起的噪声依赖,在观察数据中它可能与受保护属性不独立。
实验结果
研究问题
- RQ1在什么条件下,因果公平性模型中的独立噪声假设成立?
- RQ2现实世界数据集中的选择偏差如何破坏反事实公平性所要求的 ignorability 假设?
- RQ3在总体分布不同的选定人群中,独立噪声模型能否准确表示反事实?
- RQ4噪声依赖对路径特定公平性及其他因果公平性方法有何影响?
- RQ5为何在数据非来自随机实验时,反事实公平性在实践中无法实现?
主要发现
- 反事实公平性所需的独立噪声假设意味着 ignorability,而这种假设仅在随机对照试验中有效,且在现实世界数据中很少成立。
- 从更广泛总体中进行选择会导致噪声变量与受保护属性之间的依赖,从而违反 i.i.d. 噪声假设。
- 在美国人口普查数据集中,受雇人员中女性占比33%与总人口中50.9%的女性估计值相矛盾,从而否定了 ignorability。
- 在德国信贷数据集中,申请者中女性占比31%与总人口中51.5%的女性估计值相矛盾,进一步否定了 ignorability。
- 在法学院数据集中,女性申请者占比43.8%与总人口中49.7%的女性估计值相矛盾,提供了强有力的证据否定 ignorability。
- 即使一个预测器在真实 SCM 下是反事实公平的,由于选择引起的噪声依赖,在观察数据中它可能与受保护属性不独立。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。