[Paper Review] Collective moderation of hate, toxicity, and extremity in online discussions
This study investigates collective moderation strategies for reducing hate, toxicity, and extremity in online discussions using a four-year dataset of 130,000 German Twitter discussions during and after the 2015 migrant crisis. It finds that expressing straightforward opinions—non-factual but insult-free—most effectively reduces subsequent hate and toxicity across micro, meso, and macro levels, outperforming traditional counter-speech strategies and offering a low-barrier civic intervention.
In the digital age, hate speech poses a threat to the functioning of social media platforms as spaces for public discourse. Top-down approaches to moderate hate speech encounter difficulties due to conflicts with freedom of expression and issues of scalability. Counter speech, a form of collective moderation by citizens, has emerged as a potential remedy. Here, we aim to investigate which counter speech strategies are most effective in reducing the prevalence of hate, toxicity, and extremity on online platforms. We analyze more than 130,000 discussions on German Twitter starting at the peak of the migrant crisis in 2015 and extending over four years. We use human annotation and machine learning classifiers to identify argumentation strategies, ingroup and outgroup references, emotional tone, and different measures of discourse quality. Using matching and time-series analyses we discern the effectiveness of naturally observed counter speech strategies on the micro-level (individual tweet pairs), meso-level (entire discussions) and macro-level (over days). We find that expressing straightforward opinions, even if not factual but devoid of insults, results in the least subsequent hate, toxicity, and extremity over all levels of analyses. This strategy complements currently recommended counter speech strategies and is easy for citizens to engage in. Sarcasm can also be effective in improving discourse quality, especially in the presence of organized extreme groups. Going beyond one-shot analyses on smaller samples prevalent in most prior studies, our findings have implications for the successful management of public online spaces through collective civic moderation.
Motivation & Objective
- To investigate which counter-speech strategies are most effective in reducing hate, toxicity, and extremity in online political discussions.
- To analyze the long-term, real-world impact of citizen-led moderation on large-scale social media discourse.
- To evaluate how organized extremist groups (e.g., Reconquista Germanica and Reconquista Internet) influence the effectiveness of different discourse strategies.
- To provide evidence-based guidance for citizens and communities on how to engage in collective civic moderation without relying on top-down platform moderation.
- To move beyond small-scale, one-time experiments by analyzing discourse dynamics over four years using real-world data.
Proposed method
- Constructed a corpus of 1,150,469 tweets from German news outlets, journalists, and politicians, forming discussion trees over four years (2015–2018).
- Used human annotation and machine learning classifiers to identify argumentation strategies, outgroup/ingroup references, emotional tone, and discourse quality metrics.
- Applied matching and time-series analyses (including ARDL modeling) to assess causal effects of discourse dimensions on hate, toxicity, and extremity at micro (tweet pairs), meso (entire discussions), and macro (daily trends) levels.
- Incorporated indicators of organized extremist presence (RG and RI) to test interaction effects on counter-speech effectiveness.
- Used statistical models to isolate the impact of specific discourse strategies while controlling for confounding variables such as speaker extremity and topic.
- Published all inferred data and code for reproducibility under a material transfer agreement, with data available via OSF and code on GitHub.
Experimental results
Research questions
- RQ1Which counter-speech strategies most effectively reduce subsequent hate, toxicity, and extremity in online political discussions?
- RQ2How do the effects of different discourse strategies vary across micro-level interactions, meso-level discussion structures, and macro-level daily trends?
- RQ3How does the presence of organized extremist groups (e.g., Reconquista Germanica and Reconquista Internet) moderate the effectiveness of counter-speech strategies?
- RQ4Is expressing a straightforward opinion—regardless of factual accuracy but without insults—more effective than traditional rational or empathetic counter-speech approaches?
- RQ5Can sarcasm serve as an effective tool for improving discourse quality, particularly in the presence of coordinated extremist actors?
Key findings
- Expressing straightforward opinions—non-factual but devoid of insults—results in the lowest subsequent levels of hate, toxicity, and extremity across all levels of analysis (micro, meso, and macro).
- This strategy outperforms traditionally recommended approaches such as appeals to reason, empathy, or moral principles in reducing discourse toxicity.
- Sarcasm was found to be effective in improving discourse quality, especially when organized extremist groups were active, suggesting a potential role in countering coordinated disinformation.
- The presence of organized extremist groups like Reconquista Germanica (RG) and Reconquista Internet (RI) significantly modulates the effectiveness of counter-speech strategies, with effects varying by discourse dimension.
- The study’s findings are robust across multiple analytical levels and timeframes, demonstrating long-term, real-world impact of civic moderation strategies.
- The results provide actionable, evidence-based guidance for citizens and communities to engage in collective moderation without relying on formal platform moderation or high cognitive load.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.