[Paper Review] AHA!: Facilitating AI Impact Assessment by Generating Examples of Harms
AHA! is a generative framework that surfaces potential harms from AI deployments by filling an ethical matrix with vignettes and completing them via crowdsourcing and LLMs, then analyzing the harms across scenarios.
While demands for change and accountability for harmful AI consequences mount, foreseeing the downstream effects of deploying AI systems remains a challenging task. We developed AHA! (Anticipating Harms of AI), a generative framework to assist AI practitioners and decision-makers in anticipating potential harms and unintended consequences of AI systems prior to development or deployment. Given an AI deployment scenario, AHA! generates descriptions of possible harms for different stakeholders. To do so, AHA! systematically considers the interplay between common problematic AI behaviors as well as their potential impacts on different stakeholders, and narrates these conditions through vignettes. These vignettes are then filled in with descriptions of possible harms by prompting crowd workers and large language models. By examining 4113 harms surfaced by AHA! for five different AI deployment scenarios, we found that AHA! generates meaningful examples of harms, with different problematic AI behaviors resulting in different types of harms. Prompting both crowds and a large language model with the vignettes resulted in more diverse examples of harms than those generated by either the crowd or the model alone. To gauge AHA!'s potential practical utility, we also conducted semi-structured interviews with responsible AI professionals (N=9). Participants found AHA!'s systematic approach to surfacing harms important for ethical reflection and discovered meaningful stakeholders and harms they believed they would not have thought of otherwise. Participants, however, differed in their opinions about whether AHA! should be used upfront or as a secondary-check and noted that AHA! may shift harm anticipation from an ideation problem to a potentially demanding review problem. Drawing on our results, we discuss design implications of building tools to help practitioners envision possible harms.
Motivation & Objective
- Motivate and enable proactive anticipation of downstream harms from AI systems across diverse stakeholders.
- Develop a semi-automated process to generate ethical matrices and context-rich vignettes for harm exploration.
- Empirically evaluate harm surfaces from crowdsourcing and large language models across multiple deployment scenarios.
- Assess the added value of combining crowdsourcing with LLMs for diverse harm generation.
- Gather practitioner feedback on the utility and limitations of AHA! for responsible AI workflows.
Proposed method
- Automatically generate stakeholder sets for a deployment scenario using a large language model.
- Populate an ethical matrix where rows are stakeholders and columns are AI behaviors (16 per scenario).
- Create vignette cells describing stakeholder experiences with AI behaviors in context.
- Complete vignettes with harm descriptions using crowdsourcing and a large language model (GPT-3).
- Code and cluster harms into an eight-category taxonomy (e.g., allocational, representational, rights/agency, etc.).
- Analyze harm distributions across scenarios and behavioral dimensions using chi-square tests and qualitative coding.
Experimental results
Research questions
- RQ1Does AHA! surface meaningful harms that are relevant to each deployment scenario?
- RQ2How do varying dimensions of AI behavior (e.g., false positives vs false negatives) influence the types of harms surfaced?
- RQ3What is the comparative contribution of crowdsourcing versus an LLM in generating harms, and is a combination superior?
- RQ4What are practitioners’ views on using AHA! upfront versus as a secondary check, and what design implications arise?
- RQ5Are harms more similar within related deployment domains (e.g., hiring vs loan) than across dissimilar ones?
Key findings
- AHA! surfaces meaningful harms in 7% of crowd and GPT-3 generations, with 63.2% nonsensical among the non-meaningful portion.
- Distributions of harm categories differ significantly across deployment scenarios (Chi-square, p < .0001).
- False positives vs false negatives prime different harm types; e.g., FP leads to more allocational harms in some scenarios, FN to more representational harms.
- Crowd and GPT-3 generate comparable numbers of harms but different harm types; their combination yields more diverse harms than either source alone.
- Practitioners valued AHA! for enabling broad ethical reflection and identifying harms stakeholders might miss, but opinions varied on upfront vs secondary use and workload.
- AHA! provides design implications for tools that help practitioners envision and reflect on possible harms.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.