[论文解读] Assessing the Fairness of AI Systems: AI Practitioners' Processes, Challenges, and Needs for Support
本研究调查了AI从业者在现实环境中设计细粒度公平性评估的过程,揭示了在指标选择、利益相关者识别和数据收集方面面临的关键挑战。研究发现,组织压力——尤其是优先考虑客户影响而非边缘化群体,以及追求规模部署——显著阻碍了公平性评估的开展,呼吁在AI开发工作流程中加强利益相关者参与和结构性支持。
Various tools and practices have been developed to support practitioners in identifying, assessing, and mitigating fairness-related harms caused by AI systems. However, prior research has highlighted gaps between the intended design of these tools and practices and their use within particular contexts, including gaps caused by the role that organizational factors play in shaping fairness work. In this paper, we investigate these gaps for one such practice: disaggregated evaluations of AI systems, intended to uncover performance disparities between demographic groups. By conducting semi-structured interviews and structured workshops with thirty-three AI practitioners from ten teams at three technology companies, we identify practitioners' processes, challenges, and needs for support when designing disaggregated evaluations. We find that practitioners face challenges when choosing performance metrics, identifying the most relevant direct stakeholders and demographic groups on which to focus, and collecting datasets with which to conduct disaggregated evaluations. More generally, we identify impacts on fairness work stemming from a lack of engagement with direct stakeholders or domain experts, business imperatives that prioritize customers over marginalized groups, and the drive to deploy AI systems at scale.
研究动机与目标
- 理解AI从业者在为AI系统设计细粒度公平性评估时所面临的现实流程与挑战。
- 识别阻碍有效公平性评估的组织障碍,例如业务优先事项和缺乏利益相关者参与。
- 探索从业者在开展情境适切的公平性评估时的支持需求。
- 考察组织背景如何影响公平性评估实践,特别是在规模和部署时间线方面的关联。
提出的方法
- 在三家科技公司的10支团队中,对33名AI从业者进行了半结构化访谈。
- 组织了结构化工作坊,引导从业者通过改编自Barocas等人(2021)的公平性评估设计流程。
- 聚焦于开发多样化AI系统的团队,包括自然语言处理、欺诈检测和聊天机器人,以确保获得多样化的背景洞察。
- 收集了关于指标选择、利益相关者识别、人口群体选择以及数据收集挑战的数据。
- 通过主题编码分析发现,以识别反复出现的挑战和支持需求。
- 采用混合方法,结合访谈的定性洞察与参与式工作坊的产出。
实验结果
研究问题
- RQ1从业者在设计其AI系统的细粒度评估时,现有的流程和挑战是什么?
- RQ2从业者在设计细粒度评估时需要哪些组织支持?他们如何向领导层传达这些需求?
- RQ3组织背景如何影响从业者在评估过程、挑战和支援需求方面的表现?
主要发现
- 从业者在为细粒度公平性评估选择合适性能指标方面面临重大挑战,通常由于技术可行性与公平性目标之间的权衡。
- 识别最相关的直接利益相关者和人口群体进行评估具有困难,尤其是在利益相关者未明确定义或在开发早期未参与的情况下。
- 难以获取代表性强、高质量的代表性不足群体的数据集,是开展有意义的细粒度评估的主要障碍。
- 组织推动AI系统大规模部署的压力往往优先于彻底的公平性评估,导致评估过程仓促或流于表面。
- 缺乏与领域专家和直接利益相关者的互动,导致评估可能遗漏情境相关的公平性风险。
- 从业者通常缺乏正式的支持结构来为公平性评估发声,尤其是在评估与客户留存或快速部署等业务目标冲突时。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。