Skip to main content
QUICK REVIEW

[Paper Review] Opportunities and Risks of LLMs for Scalable Deliberation with Polis

Christopher Small, Ivan Vendrov|arXiv (Cornell University)|Jun 20, 2023
Ethics and Social Impacts of AI18 citations
TL;DR

The paper investigates how large language models (LLMs) can augment Polis for scalable deliberation, demonstrating capabilities in topic modelling, summarization, and vote prediction, while highlighting risks and mitigation strategies.

ABSTRACT

Polis is a platform that leverages machine intelligence to scale up deliberative processes. In this paper, we explore the opportunities and risks associated with applying Large Language Models (LLMs) towards challenges with facilitating, moderating and summarizing the results of Polis engagements. In particular, we demonstrate with pilot experiments using Anthropic's Claude that LLMs can indeed augment human intelligence to help more efficiently run Polis conversations. In particular, we find that summarization capabilities enable categorically new methods with immense promise to empower the public in collective meaning-making exercises. And notably, LLM context limitations have a significant impact on insight and quality of these results. However, these opportunities come with risks. We discuss some of these risks, as well as principles and techniques for characterizing and mitigating them, and the implications for other deliberative or political systems that may employ LLMs. Finally, we conclude with several open future research directions for augmenting tools like Polis with LLMs.

Motivation & Objective

  • Assess how LLMs can augment Polis to improve scalability of deliberative processes.
  • Evaluate LLM-enabled tasks such as topic modelling, summarization, moderation, and consensus discovery in Polis.
  • Identify risks (bias, hallucinations, misrepresentation) and propose mitigation strategies.
  • Demonstrate with pilot experiments using Anthropic’s Claude to augment Polis workflows.
  • Provide future directions for integrating LLMs into deliberative platforms.

Proposed method

  • Conduct pilot experiments using Anthropic Claude to operate within Polis workflows.
  • Perform topic modelling by prompting LLMs to assign topics to batches of comments.
  • Generate automated summaries and consensus statements from Polis data.
  • Evaluate vote prediction by querying LLMs about participant agreement on unseen comments.
  • Investigate the use of long context windows (8K to 100K tokens) to handle large conversations.
  • Discuss human-in-the-loop evaluation and safety mitigations for LLM outputs.

Experimental results

Research questions

  • RQ1Can LLMs reliably identify topics in Polis conversations and aid reporting?
  • RQ2To what extent can LLMs generate coherent summaries and identify group consensus from Polis data?
  • RQ3How accurately can LLMs predict votes, given prior voting histories?
  • RQ4What are the risks (bias, misinformation, moderation) of LLM-assisted Polis, and how can they be mitigated?
  • RQ5How do extended context windows impact LLM performance in large Polis conversations?

Key findings

  • Claude-generated topics aligned with manual analysis and yielded hierarchical topic structure.
  • LLMs could automatically generate concise summaries that matched manual analysis trends; however, context window limits and potential inaccuracies require careful prompting and human review.
  • A plain LLM could calibrate with high confidence on predicting whether a participant would agree or disagree with a given comment, showing strong predictive capability.
  • Using LLMs for consensus drafting surfaced example consensus statements but required live testing and governance for ethical use.
  • Topic modelling and summarization can be updated iteratively in online conversations to adapt to new comments.
  • Risks include potential misinformation, biased representations, and the need for transparent disclosure and human-in-the-loop oversight.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.