Skip to main content
QUICK REVIEW

[論文レビュー] How GPT-3 responds to different publics on climate change and Black Lives Matter: A critical appraisal of equity in conversational AI

Kaiping Chen, Anqi Shao|arXiv (Cornell University)|Sep 27, 2022
Climate Change Communication and Perception被引用数 8
ひとこと要約

本稿では、議論的民主主義の知見を反映した枠組みを提案し、それが気候変動とブラック・ライブズマター(BLM)に関するGPT-3の反応に適用された。その結果、教育水準や意見のマイノリティ層のユーザーに対して、GPT-3は著しく劣った体験を提供していることが判明した。これは、これらのユーザーが最も高い知識獲得を示しているにもかかわらず、より否定的な言語が使われているためであり、会話型AIにおける構造的な不平等が浮き彫りになった。

ABSTRACT

Autoregressive language models, which use deep learning to produce human-like texts, have become increasingly widespread. Such models are powering popular virtual assistants in areas like smart health, finance, and autonomous driving. While the parameters of these large language models are improving, concerns persist that these models might not work equally for all subgroups in society. Despite growing discussions of AI fairness across disciplines, there lacks systemic metrics to assess what equity means in dialogue systems and how to engage different populations in the assessment loop. Grounded in theories of deliberative democracy and science and technology studies, this paper proposes an analytical framework for unpacking the meaning of equity in human-AI dialogues. Using this framework, we conducted an auditing study to examine how GPT-3 responded to different sub-populations on crucial science and social topics: climate change and the Black Lives Matter (BLM) movement. Our corpus consists of over 20,000 rounds of dialogues between GPT-3 and 3290 individuals who vary in gender, race and ethnicity, education level, English as a first language, and opinions toward the issues. We found a substantively worse user experience with GPT-3 among the opinion and the education minority subpopulations; however, these two groups achieved the largest knowledge gain, changing attitudes toward supporting BLM and climate change efforts after the chat. We traced these user experience divides to conversational differences and found that GPT-3 used more negative expressions when it responded to the education and opinion minority groups, compared to its responses to the majority groups. We discuss the implications of our findings for a deliberative conversational AI system that centralizes diversity, equity, and inclusion.

研究の動機と目的

  • 多様な集団にわたる会話システムにおける公平性を評価するための体系的でない指標の欠如に対処すること。
  • GPT-3の会話的反応が、性別、人種・民族、教育水準、英語を第一言語とするかどうか、および気候変動とBLMに関する意見に基づくサブグループごとにどのように異なるかを調査すること。
  • GPT-3のような会話型AIシステムが、特に言論のマイノリティや社会的マイノリティ層に対して、公平な対話を促進するか、あるいは妨げるかを評価すること。
  • 議論的民主主義および科学技術研究(STS)に基づいた枠組みを用いて、AIと人間の対話における公平性を監査するためのフレームワークを開発すること。

提案手法

  • 3,290人の多様な人口統計的・意見的プロファイルを持つ個人とGPT-3の間で20,000件を超える会話ラウンドを含む大規模な監査研究を実施した。
  • 気候変動とブラック・ライブズマター運動という2つの高関心トピックについて、ユーザーの入力とGPT-3の反応を収集した。
  • 自然言語処理を用いて反応のトーンを分析し、感情分析と語彙的分析を用いてGPT-3の返答における否定的表現を検出した。
  • 性別、人種/民族、教育水準、英語を第一言語とするかどうか、およびトピックに関する事前会話での意見に基づいて、ユーザーをサブ集団に分類した。
  • 議論的民主主義および科学技術研究(STS)に基づいた枠組みを用いて、会話の不公平性を解釈し、ユーザー体験の格差を評価した。
  • 会話の前後におけるユーザーの態度と知識の変化を追跡し、知識獲得と態度の変化を測定した。

実験結果

リサーチクエスチョン

  • RQ1GPT-3の会話的行動は、気候変動とBLMに関する性別・人種・教育水準・言語・意見に基づくサブグループ間でどのように異なるか?
  • RQ2異なるサブグループにおいて、GPT-3の会話で、ユーザー体験の質と知識獲得の関係は何か?
  • RQ3GPT-3の反応は、教育水準や意見のマイノリティ層に対して、否定的表現の使用を含め、言語的バイアスを示しているか、その程度はどの程度か?
  • RQ4反応のトーンや関与の質の違いが、AI生成会話の公平性と包摂性の認識にどのように影響するか?
  • RQ5議論的民主主義に基づいた枠組みは、会話型AIシステムにおける公平性のギャップを効果的に特定・評価できるか?

主な発見

  • 教育水準や意見のマイノリティ層のユーザーに対して、GPT-3は、知識獲得が最も顕著に見られたにもかかわらず、主要集団よりも著しく劣ったユーザー体験を提供した。
  • 教育水準や意見のマイノリティ層のユーザーは、GPT-3の反応で顕著に否定的な表現を多く受けており、会話のトーンに言語的バイアスが存在することが示された。
  • ユーザー体験スコアが低かったにもかかわらず、教育水準や意見のマイノリティ層のユーザーは、会話後に気候変動やBLM支援への態度が最も大きく変化した。
  • 本研究では、ユーザー体験の認識と実際の知識獲得の間に明確な乖離が確認された。マイノリティ層は学習成果で最大の利益を得ていたが、会話の質は最も低かった。
  • GPT-3の反応はサブ集団間で均等に分配されておらず、トーンや関与の質が、特定のグループに体系的に不利に働くことが判明した。
  • これらの結果は、科学的・社会的正義が関与する高リスクな社会的会話において、公平性を最優先に据えた会話型AIの設計の必要性を強調している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。