[论文解读] AHA!: Facilitating AI Impact Assessment by Generating Examples of Harms
AHA!是一个生成性框架,通过用情景小品填充伦理矩阵并通过众包和大语言模型完成它们,然后在不同情景中分析潜在的有害影响。
While demands for change and accountability for harmful AI consequences mount, foreseeing the downstream effects of deploying AI systems remains a challenging task. We developed AHA! (Anticipating Harms of AI), a generative framework to assist AI practitioners and decision-makers in anticipating potential harms and unintended consequences of AI systems prior to development or deployment. Given an AI deployment scenario, AHA! generates descriptions of possible harms for different stakeholders. To do so, AHA! systematically considers the interplay between common problematic AI behaviors as well as their potential impacts on different stakeholders, and narrates these conditions through vignettes. These vignettes are then filled in with descriptions of possible harms by prompting crowd workers and large language models. By examining 4113 harms surfaced by AHA! for five different AI deployment scenarios, we found that AHA! generates meaningful examples of harms, with different problematic AI behaviors resulting in different types of harms. Prompting both crowds and a large language model with the vignettes resulted in more diverse examples of harms than those generated by either the crowd or the model alone. To gauge AHA!'s potential practical utility, we also conducted semi-structured interviews with responsible AI professionals (N=9). Participants found AHA!'s systematic approach to surfacing harms important for ethical reflection and discovered meaningful stakeholders and harms they believed they would not have thought of otherwise. Participants, however, differed in their opinions about whether AHA! should be used upfront or as a secondary-check and noted that AHA! may shift harm anticipation from an ideation problem to a potentially demanding review problem. Drawing on our results, we discuss design implications of building tools to help practitioners envision possible harms.
研究动机与目标
- 推动并实现对来自不同利益相关者的 AI 系统潜在下游有害影响的主动预判。
- 开发半自动化流程,用于生成伦理矩阵和具有丰富情境的情景小品,以探索有害影响。
- 在多种部署情景中,通过实证评估来自众包和大语言模型的有害信息面。
- 评估将众包与大语言模型结合用于多样化有害信息生成的附加价值。
- 收集从业者对 AHA! 在负责任的 AI 工作流程中的效用与局限性的反馈。
提出的方法
- 使用大语言模型自动生成一个部署情景的利益相关者集合。
- 填充伦理矩阵,其中行是利益相关者,列是 AI 行为(每个情景 16 项)。
- 创建情景小品单元,描述利益相关者在特定情境中对 AI 行为的体验。
- 使用众包和大语言模型(GPT-3)为情景小品补充有害描述。
- 将有害描述编码并聚类到八类分类法中(例如,分配性、表征、权利/主体性等)。
- 使用卡方检验和定性编码分析在情景和行为维度上的有害分布。
实验结果
研究问题
- RQ1AHA!是否在每个部署情景中揭示出相关且有意义的有害信息?
- RQ2AI 行为的不同维度(如假阳性与假阴性)如何影响揭示的有害类型?
- RQ3众包与大语言模型在生成有害信息中的相对贡献是多少,组合是否优于单一来源?
- RQ4从业者对将 AHA! 作为前置使用与作为二次检查的看法如何,以及由此产生的设计含义?
- RQ5在相关的部署领域(如招聘与贷款)内,有害信息是否比在不相关领域之间更相似?
主要发现
- AHA!在众包和 GPT-3 生成中揭示出有意义的有害信息占比 7%,其中在非有意义部分中有 63.2% 是无意义的。
- 不同部署情景之间的有害类别分布存在显著差异(卡方检验,p < .0001)。
- 假阳性与假阴性倾向不同的有害类型;例如,在某些情景中,假阳性导致更多的分配性有害,假阴性导致更多的表征性有害。
- 众包与 GPT-3 产生的有害数量相当,但类型不同;它们的组合比任一来源单独使用产生更丰富的有害类型。
- 从业者认为 AHA!有助于促进广泛的伦理反思并识别利益相关者可能忽略的有害,但对于前置使用与二次使用以及工作量的观点存在差异。
- AHA!为帮助从业者设想并反思可能有的有害影响的工具提供了设计含义。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。