[论文解读] Fairness for Image Generation with Uncertain Sensitive Attributes
本文提出了条件比例代表性(CPR),一种图像生成中的公平性定义,确保重建图像按比例反映受保护群体,而无需预先定义群体身份。它证明了通过Langevin动力学进行后验抽样可无意识地实现CPR——无需依赖群体标签,从而对种族等模糊或有争议的敏感属性具有鲁棒性。
This work tackles the issue of fairness in the context of generative procedures, such as image super-resolution, which entail different definitions from the standard classification setting. Moreover, while traditional group fairness definitions are typically defined with respect to specified protected groups -- camouflaging the fact that these groupings are artificial and carry historical and political motivations -- we emphasize that there are no ground truth identities. For instance, should South and East Asians be viewed as a single group or separate groups? Should we consider one race as a whole or further split by gender? Choosing which groups are valid and who belongs in them is an impossible dilemma and being "fair" with respect to Asians may require being "unfair" with respect to South Asians. This motivates the introduction of definitions that allow algorithms to be \emph{oblivious} to the relevant groupings. We define several intuitive notions of group fairness and study their incompatibilities and trade-offs. We show that the natural extension of demographic parity is strongly dependent on the grouping, and \emph{impossible} to achieve obliviously. On the other hand, the conceptually new definition we introduce, Conditional Proportional Representation, can be achieved obliviously through Posterior Sampling. Our experiments validate our theoretical results and achieve fair image reconstruction using state-of-the-art generative models.
研究动机与目标
- 解决传统群体公平性定义在种族等模糊或有争议的敏感属性下失效的图像生成中的公平性问题。
- 认识到定义受保护群体(如亚洲人 vs. 南亚人)本质上具有政治性和情境依赖性,从而削弱了标准公平性定义的有效性。
- 提出一种不依赖预定义分组的公平性框架,实现对所有可能分组的公平性。
- 开发并验证一种无需访问敏感属性标签即可实现公平性的方法,确保对群体模糊性的鲁棒性。
- 证明现有方法如PULSE即使在多样化数据上训练,仍因模型先验和数据不平衡而加剧图像重建中的偏见。
提出的方法
- 提出条件比例代表性(CPR),一种公平性概念,确保从某一群体生成样本的概率与其在数据中的先验出现频率成比例。
- 证明CPR可通过后验抽样结合Langevin动力学无意识地实现——无需知晓群体身份。
- 形式化CPR与代表性人口均等性(RDP)之间的不相容性,表明RDP对任意群体定义具有强依赖性。
- 使用贝叶斯推断建模图像重建的后验分布,通过随机抽样实现公平性。
- 使用深度生成模型和Langevin动力学实现该方法,从后验中抽样,避免确定性重建带来的偏差。
- 在FFHQ和AFHQ数据集上的图像超分辨率任务中,通过实证验证该方法,与PULSE及其他基线方法进行比较。
实验结果
研究问题
- RQ1是否可以定义一种与种族或民族等任意或有争议的群体定义无关的图像生成公平性?
- RQ2哪些公平性定义与无意识性兼容——即在无需群体标签的情况下,对所有可能分组实现公平性?
- RQ3为何现有生成模型如PULSE即使在多样化数据上训练,仍会产生种族偏见的重建结果?
- RQ4能否通过单一算法同时实现对所有可能分组的公平性,而无需事先知晓敏感属性?
- RQ5是否存在一种图像生成中的公平性概念,可确保群体的按比例代表性,而无需依赖人口均等性或独立性约束?
主要发现
- 代表性人口均等性(RDP)对受保护群体的选择具有强依赖性,且无法无意识地实现,因此不适用于种族等模糊属性。
- 条件比例代表性(CPR)是一种概念上全新的公平性定义,可在无需群体标签的情况下确保群体的按比例代表性。
- 通过Langevin动力学进行后验抽样是目前唯一可完全无意识地实现CPR的方法,因此对群体模糊性具有鲁棒性。
- 实验表明,当狗占80%时,PULSE将80%的猫重建为狗,显示出由于模型先验和数据不平衡导致的严重偏见。
- 在平衡设置中,后验抽样满足RDP、SPE和PR,而PULSE无法维持公平性,验证了理论结论。
- 即使群体代表性公平性得到满足,代表性群体的重建质量差仍会破坏公平性——强调了需要具备质量意识的公平性度量。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。