Skip to main content
QUICK REVIEW

[Paper Review] How GPT-3 responds to different publics on climate change and Black Lives Matter: A critical appraisal of equity in conversational AI

Kaiping Chen, Anqi Shao|arXiv (Cornell University)|Sep 27, 2022
Climate Change Communication and Perception8 citations
TL;DR

This paper proposes a deliberative democracy-informed framework to assess equity in human-AI dialogues, applying it to GPT-3's responses on climate change and Black Lives Matter. It finds that GPT-3 delivers significantly worse user experiences to education and opinion minority groups—despite these users showing the largest knowledge gains—due to more negative language in responses, highlighting systemic inequities in conversational AI.

ABSTRACT

Autoregressive language models, which use deep learning to produce human-like texts, have become increasingly widespread. Such models are powering popular virtual assistants in areas like smart health, finance, and autonomous driving. While the parameters of these large language models are improving, concerns persist that these models might not work equally for all subgroups in society. Despite growing discussions of AI fairness across disciplines, there lacks systemic metrics to assess what equity means in dialogue systems and how to engage different populations in the assessment loop. Grounded in theories of deliberative democracy and science and technology studies, this paper proposes an analytical framework for unpacking the meaning of equity in human-AI dialogues. Using this framework, we conducted an auditing study to examine how GPT-3 responded to different sub-populations on crucial science and social topics: climate change and the Black Lives Matter (BLM) movement. Our corpus consists of over 20,000 rounds of dialogues between GPT-3 and 3290 individuals who vary in gender, race and ethnicity, education level, English as a first language, and opinions toward the issues. We found a substantively worse user experience with GPT-3 among the opinion and the education minority subpopulations; however, these two groups achieved the largest knowledge gain, changing attitudes toward supporting BLM and climate change efforts after the chat. We traced these user experience divides to conversational differences and found that GPT-3 used more negative expressions when it responded to the education and opinion minority groups, compared to its responses to the majority groups. We discuss the implications of our findings for a deliberative conversational AI system that centralizes diversity, equity, and inclusion.

Motivation & Objective

  • To address the lack of systematic metrics for assessing equity in dialogue systems across diverse populations.
  • To investigate how GPT-3's conversational responses vary across subgroups defined by gender, race, education, language, and opinion on climate change and BLM.
  • To evaluate whether conversational AI systems like GPT-3 promote or hinder equitable dialogue, especially for marginalized or minority opinion groups.
  • To develop a framework grounded in deliberative democracy and science and technology studies for auditing equity in AI-human interactions.

Proposed method

  • Conducted a large-scale auditing study with over 20,000 dialogue rounds between GPT-3 and 3,290 individuals representing diverse demographics and opinion profiles.
  • Collected user inputs and GPT-3 responses on two high-salience topics: climate change and the Black Lives Matter movement.
  • Applied natural language processing to analyze response tone, using sentiment and lexical analysis to detect negative expressions in GPT-3’s replies.
  • Classified users into subpopulations based on gender, race/ethnicity, education level, English as a first language, and pre-chat opinion on the topics.
  • Used a framework informed by deliberative democracy and science and technology studies to interpret conversational inequities and assess user experience disparities.
  • Tracked changes in user attitudes and knowledge before and after dialogue to measure knowledge gain and attitude shift.

Experimental results

Research questions

  • RQ1How does GPT-3’s conversational behavior differ across demographic and opinion-based subpopulations on climate change and Black Lives Matter?
  • RQ2What is the relationship between user experience quality and knowledge gain in GPT-3 dialogues across different subgroups?
  • RQ3To what extent do GPT-3 responses exhibit linguistic bias, particularly in the use of negative expressions, toward education and opinion minority groups?
  • RQ4How do differences in response tone and engagement quality affect the perceived fairness and inclusivity of AI-generated dialogue?
  • RQ5Can a deliberative democracy-informed framework effectively identify and assess equity gaps in conversational AI systems?

Key findings

  • GPT-3 delivered a substantively worse user experience to education and opinion minority groups compared to majority groups, despite these users showing the largest knowledge gains.
  • Users from education and opinion minority backgrounds received significantly more negative expressions in GPT-3’s responses, indicating linguistic bias in conversational tone.
  • Despite lower user experience scores, education and opinion minority groups exhibited the greatest attitude shifts toward supporting climate change and Black Lives Matter efforts after dialogue.
  • The study identified a clear disconnect between perceived user experience and actual knowledge acquisition, with marginalized groups benefiting most in learning outcomes but least in conversational quality.
  • The findings reveal that GPT-3’s responses are not equitably distributed across subpopulations, with tone and engagement quality systematically disadvantaging certain groups.
  • The results underscore the need for equity-centered design in conversational AI, particularly in high-stakes societal dialogues involving science and social justice.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.