[论文解读] Group Retention when Using Machine Learning in Sequential Decision Making: the Interplay between User Dynamics and Fairness
本文研究了机器学习中的公平性准则如何影响序列决策系统中长期的群体代表性,表明不匹配的公平性定义会随时间加剧代表性差异。本文提出了一个用户动态模型,以使公平性准则与留存驱动因素对齐,证明只有与用户离开/到达机制相匹配的公平性准则才能维持群体代表性,实证结果证实EqLos公平性可维持群体平衡,而其他准则则导致边缘化。
Machine Learning (ML) models trained on data from multiple demographic groups can inherit representation disparity (Hashimoto et al., 2018) that may exist in the data: the model may be less favorable to groups contributing less to the training process; this in turn can degrade population retention in these groups over time, and exacerbate representation disparity in the long run. In this study, we seek to understand the interplay between ML decisions and the underlying group representation, how they evolve in a sequential framework, and how the use of fairness criteria plays a role in this process. We show that the representation disparity can easily worsen over time under a natural user dynamics (arrival and departure) model when decisions are made based on a commonly used objective and fairness criteria, resulting in some groups diminishing entirely from the sample pool in the long run. It highlights the fact that fairness criteria have to be defined while taking into consideration the impact of decisions on user dynamics. Toward this end, we explain how a proper fairness criterion can be selected based on a general user dynamics model.
研究动机与目标
- 理解公平性准则对序列机器学习系统中群体代表性之长期影响。
- 建模用户动态(到达与离开)如何与机器学习决策及公平性干预相互作用。
- 识别公平性准则在何种条件下会加剧或缓解随时间推移的代表性差异。
- 提出一个框架,用于选择与用户留存动态对齐的公平性准则,以防止群体数量减少。
提出的方法
- 构建一个包含两个人口群体的序列决策模型,追踪随时间变化的特征分布和群体规模。
- 引入一个用户留存模型,其中离开/到达率取决于决策结果(例如,假阳率/假阴率、感知损失)。
- 分析在不同用户动态下,四种公平性准则(统计独立性、平等机会、平等机会率、EqLos)的影响。
- 利用通用用户动态模型,推导出公平性准则维持或加剧群体代表性的条件。
- 通过模拟实验,改变特征分布和到达率,评估长期群体比例的演变。
- 将贝叶斯最优决策与从系统中剩余用户学习到的决策进行比较,以评估模型准确度下降的影响。
实验结果
研究问题
- RQ1在序列决策系统中,常用公平性准则如何随时间影响群体代表性?
- RQ2在何种条件下,公平性干预会导致弱势群体被边缘化甚至灭绝?
- RQ3为何一些公平性准则尽管满足统计公平性,仍无法维持群体代表性?
- RQ4公平性准则与用户留存机制之间的对齐程度如何影响长期系统公平性?
- RQ5用户动态模型能否指导选择可防止代表性差异恶化的公平性准则?
主要发现
- 在常用的公平性准则(如统计独立性和平等机会)下,弱势群体可能随时间被边缘化甚至从系统中消失,尤其是在其损失低于多数群体时。
- EqLos公平性准则——即在各群体间均衡损失——即使在到达率不对称的情况下,也能在无限时间范围内维持群体代表性。
- 当公平性准则未与用户离开的关键因素(如假阴性率或感知损失)对齐时,无论是否具备公平性保证,代表性差异都会加剧。
- 从不断缩小的用户群体中学习决策会导致模型准确度下降,并加速代表性差异的恶化,尤其在Simple、EqOpt和StatPar公平性下更为显著。
- 在用户动态由子群体特定感知损失驱动的系统中,没有任何公平性准则能维持群体代表性,凸显了机制感知公平性的重要性。
- 实证结果证实,当决策从不断减少的用户基础中学习时,群体代表性差异最为严重,尤其是在公平性定义不匹配的情况下。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。