[Paper Review] Embedding Democratic Values into Social Media AIs via Societal Objective Functions
This paper proposes a method to embed democratic values—specifically, reducing partisan animosity—into social media AI feed ranking algorithms by translating established social science constructs into societal objective functions. Using large language models (LLMs) to score posts for anti-democratic attitudes, the authors demonstrate in three studies that downranking such content significantly reduces partisan animosity without harming user engagement.
Can we design artificial intelligence (AI) systems that rank our social media feeds to consider democratic values such as mitigating partisan animosity as part of their objective functions? We introduce a method for translating established, vetted social scientific constructs into AI objective functions, which we term societal objective functions, and demonstrate the method with application to the political science construct of anti-democratic attitudes. Traditionally, we have lacked observable outcomes to use to train such models, however, the social sciences have developed survey instruments and qualitative codebooks for these constructs, and their precision facilitates translation into detailed prompts for large language models. We apply this method to create a democratic attitude model that estimates the extent to which a social media post promotes anti-democratic attitudes, and test this democratic attitude model across three studies. In Study 1, we first test the attitudinal and behavioral effectiveness of the intervention among US partisans (N=1,380) by manually annotating (alpha=.895) social media posts with anti-democratic attitude scores and testing several feed ranking conditions based on these scores. Removal (d=.20) and downranking feeds (d=.25) reduced participants' partisan animosity without compromising their experience and engagement. In Study 2, we scale up the manual labels by creating the democratic attitude model, finding strong agreement with manual labels (rho=.75). Finally, in Study 3, we replicate Study 1 using the democratic attitude model instead of manual labels to test its attitudinal and behavioral impact (N=558), and again find that the feed downranking using the societal objective function reduced partisan animosity (d=.25). This method presents a novel strategy to draw on social science theory and methods to mitigate societal harms in social media AIs.
Motivation & Objective
- To address the growing societal harm caused by social media AI algorithms that amplify partisan animosity and undermine democratic discourse.
- To overcome the lack of observable, algorithmically tractable outcomes for training AI systems to mitigate partisan animosity.
- To develop a method that translates rigorously validated social science constructs—like anti-democratic attitudes—into actionable, measurable objectives for AI systems.
- To evaluate whether integrating these societal objective functions into feed ranking algorithms can reduce partisan animosity while preserving user engagement.
- To ensure ethical deployment by addressing risks related to freedom of expression, value trade-offs, and disparate impacts on marginalized communities.
Proposed method
- Translating established social science survey instruments and qualitative codebooks for anti-democratic attitudes into detailed, LLM-interpretable prompts.
- Training a large language model (LLM) on these prompts to create a 'democratic attitude model' that estimates the extent to which a social media post promotes anti-democratic attitudes.
- Using the democratic attitude model to generate automated labels for social media posts, validated against manual annotations (weighted Kappa = .895).
- Implementing a societal objective function in feed ranking that downranks or removes posts scoring high on anti-democratic attitudes.
- Designing controlled experiments with randomized feed ranking conditions to test the impact on user attitudes and behaviors.
- Validating the method across three studies: manual annotation (Study 1), model scaling (Study 2), and automated intervention (Study 3).
Experimental results
Research questions
- RQ1Can social media feed ranking algorithms that incorporate a societal objective function based on anti-democratic attitudes reduce partisan animosity among users?
- RQ2Does downranking or removing posts with high anti-democratic attitude scores affect user engagement and perceived experience?
- RQ3How well does an LLM-based democratic attitude model align with human-annotated scores of anti-democratic content?
- RQ4Can the societal objective function method be scaled and replicated in real-world settings without compromising user experience?
- RQ5What are the ethical risks and trade-offs of encoding societal values like democratic norms into algorithmic ranking systems?
Key findings
- Removal of high-anti-democratic-content posts reduced partisan animosity by a small-to-moderate effect size (d = .20) in Study 1.
- Downranking such content reduced partisan animosity by a moderate effect size (d = .25) in Study 1, with no significant negative impact on user engagement.
- The LLM-based democratic attitude model showed strong agreement with manual annotations (Spearman’s rho = .75), validating its reliability.
- In Study 3, replicating the downranking intervention using the automated model instead of manual labels still reduced partisan animosity (d = .25), confirming scalability.
- The method successfully translated social science constructs into algorithmic objectives, enabling measurable mitigation of societal harms in AI systems.
- The approach maintains user engagement while reducing negative political affect, suggesting a viable path to value-aligned AI design in social media platforms.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.