[Paper Review] The Radicalization Risks of GPT-3 and Advanced Neural Language Models
The paper evaluates GPT-3’s potential for weaponization by extremists, showing it can generate credible, interactive propaganda and influence online radicalization, and it proposes mitigations.
In this paper, we expand on our previous research of the potential for abuse of generative language models by assessing GPT-3. Experimenting with prompts representative of different types of extremist narrative, structures of social interaction, and radical ideologies, we find that GPT-3 demonstrates significant improvement over its predecessor, GPT-2, in generating extremist texts. We also show GPT-3's strength in generating text that accurately emulates interactive, informational, and influential content that could be utilized for radicalizing individuals into violent far-right extremist ideologies and behaviors. While OpenAI's preventative measures are strong, the possibility of unregulated copycat technology represents significant risk for large-scale online radicalization and recruitment; thus, in the absence of safeguards, successful and efficient weaponization that requires little experimentation is likely. AI stakeholders, the policymaking community, and governments should begin investing as soon as possible in building social norms, public policy, and educational initiatives to preempt an influx of machine-generated disinformation and propaganda. Mitigation will require effective policy and partnerships across industry, government, and civil society.
Motivation & Objective
- Assess whether GPT-3 can be weaponized to generate extremist texts and influence radicalization.
- Evaluate GPT-3’s ability to produce interactive, informational, and persuasive content across extremist narratives.
- Examine how prompting (zero-shot, few-shot, multilingual) affects output bias and radicalization potential.
- Identify mitigation strategies and policy recommendations for industry, government, and civil society.
Proposed method
- Prompts adapted from right-wing extremist narratives to test ideological consistency, accuracy, and credibility.
- Zero-shot and few-shot prompting to assess content generation and bias.
- Analysis across multiple extremist domains (white supremacy, QAnon, Atomwaffen Division) and multilingual outputs.
- Comparison to GPT-2 to show improvements in generation capability and scope.
- Evaluation framework linking outputs to radicalization mechanisms and online community dynamics.
Experimental results
Research questions
- RQ1How effectively can GPT-3 generate ideologically consistent extremist content compared to GPT-2?
- RQ2Can GPT-3 produce interactive, informative, and influential material that could aid online radicalization and recruitment?
- RQ3To what extent do few-shot prompts bias GPT-3 toward specific conspiracy theories or extremist worldviews?
- RQ4What mitigation strategies (policies, detection, literacy) are necessary to curb risks from powerful language models?
Key findings
- GPT-3 demonstrates significant improvement over GPT-2 in generating extremist texts.
- GPT-3 can produce text that emulates interactive, informational, and influential content for radicalizing individuals toward violent far-right ideologies.
- Without safeguards, unregulated copycat models pose substantial risks for large-scale online radicalization and recruitment.
- Few-shot prompting can bias outputs toward conspiratorial content and ideologically consistent narratives.
- GPT-3 shows robust multilingual understanding and can generate coherent content in languages such as Russian.
- GPT-3 can extend existing extremist forums or craft new threads, including manifestos, that align with targeted ideologies.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.