Skip to main content
QUICK REVIEW

[论文解读] Opportunities and Risks of LLMs for Scalable Deliberation with Polis

Christopher Small, Ivan Vendrov|arXiv (Cornell University)|Jun 20, 2023
Ethics and Social Impacts of AI被引用 18
一句话总结

本文研究大语言模型(LLMs)如何增强 Polis 以实现可扩展的商议,展示在主题建模、摘要和投票预测方面的能力,同时强调风险及缓解策略。

ABSTRACT

Polis is a platform that leverages machine intelligence to scale up deliberative processes. In this paper, we explore the opportunities and risks associated with applying Large Language Models (LLMs) towards challenges with facilitating, moderating and summarizing the results of Polis engagements. In particular, we demonstrate with pilot experiments using Anthropic's Claude that LLMs can indeed augment human intelligence to help more efficiently run Polis conversations. In particular, we find that summarization capabilities enable categorically new methods with immense promise to empower the public in collective meaning-making exercises. And notably, LLM context limitations have a significant impact on insight and quality of these results. However, these opportunities come with risks. We discuss some of these risks, as well as principles and techniques for characterizing and mitigating them, and the implications for other deliberative or political systems that may employ LLMs. Finally, we conclude with several open future research directions for augmenting tools like Polis with LLMs.

研究动机与目标

  • 评估 LLMs 如何增强 Polis 以提升商议过程的可扩展性。
  • 评估在 Polis 中由 LLM 支持的任务,如主题建模、摘要、审核和共识发现。
  • 识别风险(偏见、幻觉、错误表述)并提出缓解策略。
  • 使用 Anthropic 的 Claude 进行试点实验以增强 Polis 工作流进行演示。
  • 提供将 LLMs 整合到商议平台中的未来方向。

提出的方法

  • 进行试点实验,使用 Anthropic Claude 在 Polis 工作流中运行。
  • 通过提示 LLM 将主题分配给一批批评论来进行主题建模。
  • 从 Polis 数据生成自动摘要和共识陈述。
  • 通过向 LLM 询问对未见评论的参与者同意情况来评估投票预测。
  • 研究使用长上下文窗口(8K 至 100K token)以处理大型对话。
  • 讨论人机循环评估以及对 LLM 输出的安全缓解措施。

实验结果

研究问题

  • RQ1LLMs 能否可靠地在 Polis 对话中识别主题并帮助报道?
  • RQ2LLMs 在多大程度上能够生成连贯的摘要并从 Polis 数据中识别集团共识?
  • RQ3在既有投票历史的前提下,LLMs 能多准确地预测投票?
  • RQ4LLM 辅助的 Polis 存在哪些风险(偏见、错误信息、审核)?以及如何缓解?
  • RQ5扩展上下文窗口如何影响大型 Polis 对话中的 LLM 性能?

主要发现

  • Claude 生成的主题与人工分析一致,并产生了分层主题结构。
  • LLMs 可以自动生成与人工分析趋势一致的简明摘要;然而,上下文窗口限制和潜在不准确性需要谨慎提示和人工审核。
  • 普通 LLM 在预测参与者是否会同意或不同意给定评论方面可以以很高的置信度进行校准,显示出强大的预测能力。
  • 使用 LLMs 进行共识起草时浮现出示例共识陈述,但需要实时测试和治理以实现伦理使用。
  • 主题建模和摘要可以在在线对话中迭代更新,以适应新评论。
  • 风险包括潜在错误信息、偏见表现,以及需要透明披露和人机循环监督。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。