[论文解读] Embedding Democratic Values into Social Media AIs via Societal Objective Functions
本文提出一种方法,通过将既有的社会科学概念转化为社会性目标函数,将民主价值观——特别是减少党派敌意——嵌入社交媒体AI推荐算法中。利用大语言模型(LLMs)对帖子进行反民主态度评分,作者在三个研究中证明,降低此类内容的推荐排名能显著减少党派敌意,且不会损害用户参与度。
Can we design artificial intelligence (AI) systems that rank our social media feeds to consider democratic values such as mitigating partisan animosity as part of their objective functions? We introduce a method for translating established, vetted social scientific constructs into AI objective functions, which we term societal objective functions, and demonstrate the method with application to the political science construct of anti-democratic attitudes. Traditionally, we have lacked observable outcomes to use to train such models, however, the social sciences have developed survey instruments and qualitative codebooks for these constructs, and their precision facilitates translation into detailed prompts for large language models. We apply this method to create a democratic attitude model that estimates the extent to which a social media post promotes anti-democratic attitudes, and test this democratic attitude model across three studies. In Study 1, we first test the attitudinal and behavioral effectiveness of the intervention among US partisans (N=1,380) by manually annotating (alpha=.895) social media posts with anti-democratic attitude scores and testing several feed ranking conditions based on these scores. Removal (d=.20) and downranking feeds (d=.25) reduced participants' partisan animosity without compromising their experience and engagement. In Study 2, we scale up the manual labels by creating the democratic attitude model, finding strong agreement with manual labels (rho=.75). Finally, in Study 3, we replicate Study 1 using the democratic attitude model instead of manual labels to test its attitudinal and behavioral impact (N=558), and again find that the feed downranking using the societal objective function reduced partisan animosity (d=.25). This method presents a novel strategy to draw on social science theory and methods to mitigate societal harms in social media AIs.
研究动机与目标
- 解决社交媒体AI算法加剧党派敌意、损害民主话语所造成的日益严重的社会危害。
- 克服训练AI系统以缓解党派敌意时缺乏可观测、可算法处理的结果这一难题。
- 开发一种将经过严格验证的社会科学概念(如反民主态度)转化为AI系统可执行、可度量目标的方法。
- 评估在推荐算法中整合这些社会性目标函数是否能在维持用户参与度的同时减少党派敌意。
- 通过解决言论自由、价值权衡以及对边缘化群体的非均衡影响等相关风险,确保伦理化部署。
提出的方法
- 将既有的社会科学调查工具和反民主态度的定性编码手册转化为详细、可被大语言模型理解的提示。
- 基于这些提示训练大语言模型(LLM),构建一个‘民主态度模型’,用以估算社交媒体帖子在多大程度上宣扬反民主态度。
- 利用民主态度模型为社交媒体帖子生成自动化标签,并与人工标注结果进行验证(加权 kappa = .895)。
- 在推荐算法中实施社会性目标函数,对反民主态度评分较高的内容进行降权或移除。
- 设计受控实验,采用随机化的推荐算法条件,测试其对用户态度和行为的影响。
- 在三个研究中验证该方法:人工标注(研究1)、模型扩展(研究2)和自动化干预(研究3)。
实验结果
研究问题
- RQ1基于反民主态度的社会性目标函数的社交媒体推荐算法,能否减少用户之间的党派敌意?
- RQ2对高反民主态度评分内容进行降权或移除,是否会影响用户参与度和感知体验?
- RQ3基于LLM的民主态度模型与人工标注的反民主内容评分之间的吻合度如何?
- RQ4该社会性目标函数方法是否可在现实环境中实现规模化和可复现,同时不损害用户体验?
- RQ5在算法推荐系统中编码诸如民主规范等社会价值观,存在哪些伦理风险和权衡?
主要发现
- 在研究1中,移除高反民主态度内容使党派敌意降低了小到中等程度的效应量(d = .20)。
- 在研究1中,对高反民主态度内容进行降权处理使党派敌意降低了中等程度的效应量(d = .25),且对用户参与度无显著负面影响。
- 基于LLM的民主态度模型与人工标注结果高度一致(Spearman等级相关系数 = .75),验证了其可靠性。
- 在研究3中,使用自动化模型替代人工标签实施降权干预,同样显著降低了党派敌意(d = .25),证实了该方法的可扩展性。
- 该方法成功地将社会科学概念转化为算法目标,使AI系统能够可度量地缓解社会性危害。
- 该方法在降低负面政治情绪的同时维持了用户参与度,表明其为社交媒体平台实现价值对齐的AI设计提供了切实可行的路径。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。