[Paper Review] Protecting Society from AI Misuse: When are Restrictions on Capabilities Warranted?
The paper argues for targeted restrictions on AI capabilities and some non-AI capabilities to prevent misuse, developing a framework and taxonomy to assess when such interventions are warranted.
Artificial intelligence (AI) systems will increasingly be used to cause harm as they grow more capable. In fact, AI systems are already starting to be used to automate fraudulent activities, violate human rights, create harmful fake images, and identify dangerous toxins. To prevent some misuses of AI, we argue that targeted interventions on certain capabilities will be warranted. These restrictions may include controlling who can access certain types of AI models, what they can be used for, whether outputs are filtered or can be traced back to their user, and the resources needed to develop them. We also contend that some restrictions on non-AI capabilities needed to cause harm will be required. Though capability restrictions risk reducing use more than misuse (facing an unfavorable Misuse-Use Tradeoff), we argue that interventions on capabilities are warranted when other interventions are insufficient, the potential harm from misuse is high, and there are targeted ways to intervene on capabilities. We provide a taxonomy of interventions that can reduce AI misuse, focusing on the specific steps required for a misuse to cause harm (the Misuse Chain), and a framework to determine if an intervention is warranted. We apply this reasoning to three examples: predicting novel toxins, creating harmful images, and automating spear phishing campaigns.
Motivation & Objective
- Motivate that AI misuses will escalate as capabilities grow and that targeted capability restrictions can reduce harm.
- Propose a taxonomy of interventions that limit misuse while accounting for costs and tradeoffs.
- Develop a framework (Misuse Chain) to determine when restricting capabilities is warranted.
- Argue that restrictions can extend beyond AI capabilities to include non-AI capability controls.
- Apply the framework to concrete misuse scenarios to illustrate practical guidance.
Proposed method
- Develop a taxonomy of interventions that can reduce AI misuse by targeting specific steps in the misuse process.
- Introduce the Misuse Chain framework to map where interventions can disrupt harm.
- Discuss criteria for when capability restrictions are warranted (risk, sufficiency of other interventions, targeted feasibility).
- Provide examples illustrating the application to three misuse domains: novel toxin prediction, harmful image creation, and automated spear phishing.
- Contrast capability restrictions with the Misuse-Use tradeoff to justify selective deployment.
Experimental results
Research questions
- RQ1Under what conditions are restrictions on AI capabilities warranted to prevent misuse?
- RQ2What forms can capability restrictions take (e.g., access, activities, outputs, traceability, resources) and how do they interact with non-AI restrictions?
- RQ3How can a systematic framework (Misuse Chain) assess the effectiveness and necessity of interventions across misuse scenarios?
- RQ4What are concrete examples where targeted capability restrictions could reduce harm without unduly hindering beneficial use?
Key findings
- Targeted restrictions on capabilities can be warranted when other interventions are insufficient and potential harm from misuse is high.
- Restrictions may include access controls, allowed use cases, output filtering or traceability, and resource limitations.
- Restrictions on non-AI capabilities needed to cause harm may also be required.
- A Misuse Chain framework helps identify intervention points that can meaningfully reduce harm in a structured way.
- Applied to toxins prediction, harmful image generation, and spear phishing, illustrating how interventions can disrupt the misuse process.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.