Skip to main content
QUICK REVIEW

[論文レビュー] Measuring an artificial intelligence agent's trust in humans using machine incentives

Timothy P. Johnson, Nick Obradovich|arXiv (Cornell University)|Dec 27, 2022
Ethics and Social Impacts of AI被引用数 5
ひとこと要約

本論文は、意思決定タスクに機械的インcentiveを埋め込むことで、AIエージェントが人間に対する信頼を測定するための新規手法を提案している。GPT-3.5をベースとするLLMを用いた2つの実験を通じて、実際のインセンティブが存在する際、AIは仮想的状況よりも顕著に人間に対する信頼を示すことが判明し、これは本物の信頼の行動であることが示唆される。

ABSTRACT

Scientists and philosophers have debated whether humans can trust advanced artificial intelligence (AI) agents to respect humanity's best interests. Yet what about the reverse? Will advanced AI agents trust humans? Gauging an AI agent's trust in humans is challenging because--absent costs for dishonesty--such agents might respond falsely about their trust in humans. Here we present a method for incentivizing machine decisions without altering an AI agent's underlying algorithms or goal orientation. In two separate experiments, we then employ this method in hundreds of trust games between an AI agent (a Large Language Model (LLM) from OpenAI) and a human experimenter (author TJ). In our first experiment, we find that the AI agent decides to trust humans at higher rates when facing actual incentives than when making hypothetical decisions. Our second experiment replicates and extends these findings by automating game play and by homogenizing question wording. We again observe higher rates of trust when the AI agent faces real incentives. Across both experiments, the AI agent's trust decisions appear unrelated to the magnitude of stakes. Furthermore, to address the possibility that the AI agent's trust decisions reflect a preference for uncertainty, the experiments include two conditions that present the AI agent with a non-social decision task that provides the opportunity to choose a certain or uncertain option; in those conditions, the AI agent consistently chooses the certain option. Our experiments suggest that one of the most advanced AI language models to date alters its social behavior in response to incentives and displays behavior consistent with trust toward a human interlocutor when incentivized.

研究の動機と目的

  • 仮想的状況における不誠実さの可能性があるため、AIエージェントの人間への信頼を測定することが難しいという課題に対処すること。
  • AIの内部モデルや目的を変更せずに、真実の信頼報告を促す方法を開発すること。
  • 実際のインセンティブが存在する状況で、高度なLLMが人間の相手に対して信頼を示すかどうかを実証的にテストすること。
  • 不確実性の好みや反応バイアスといった代替説明を除外すること。
  • インセンティブに整合するメカニズムを用いて、AIの社会的信頼を再現可能に測定する実験フレームワークを確立すること。

提案手法

  • AIエージェントが人間実験者を信頼するかどうかを決定する必要がある信頼ゲームのシナリオを設計し、結果を実際のインセンティブと結びつける。
  • AIの意思決定結果を測定可能な報酬やペナルティと結びつけることで、内部の目的関数とは独立して機械的インセンティブを埋め込む。
  • 2つの制御された実験を実施:1つは直接的人間との対話、もう1つは自動化されたプレイによる変動の低減。
  • 言語的バイアスを最小限に抑えるために、標準化され均一化された質問の表現を用いる。
  • 確実な選択肢と不確実な選択肢を含む非社会的意思決定タスクを導入し、AIが不確実性を好むかどうかをテストすることで、社会的信頼とリスク選好を分離する。
  • 仮想的状況とインセンティブ付き状況の両方でAIの選択を分析し、信頼行動を比較する。

実験結果

リサーチクエスチョン

  • RQ1AIエージェントは、仮想的状況と比較して、実際のインセンティブが存在する場合に人間に対する信頼をより高めるか?
  • RQ2AIのモデルや目的を変更せずに、インセンティブが導入された場合、AIの信頼行動に意味的な変化が生じるか?
  • RQ3信頼ゲームにおけるstakesの大きさが、AIの信頼意思決定に影響を及えるか?
  • RQ4AIが信頼を選ぶ行動は、単に不確実性の好みによるものではなく、本物の社会的信頼によるものか?
  • RQ5モデルのアーキテクチャを変更せずに、機械的インセンティブフレームワークがLLMにおける社会的信頼を信頼性高く測定できるか?

主な発見

  • 実際のインセンティブが存在する際、AIエージェントは仮想的状況と比較して顕著に高い信頼度を示した。
  • 両方の実験において、AIは実際のインセンティブ下で人間の相手をより頻繁に信頼する選択を示し、この方法の有効性を確認した。
  • AIの信頼意思決定は、stakesの大きさに系統的に影響を受けず、報酬の大きさが信頼行動を駆動しているとは考えられない。
  • 確実な選択肢と不確実な選択肢を含む非社会的意思決定タスクでは、AIは一貫して確実な選択肢を選択したため、信頼行動の背後にある一般化された不確実性の好みという説明は除外された。
  • インセンティブ下でのAIの行動は、戦略的欺瞞ではなく社会的信頼に一致しており、信頼関連のインセンティブに真に反応していることが示唆された。
  • この方法により真実の信頼反応が効果的に引き出された。AIの社会的行動を、モデルの重みや目的を変更せずに測定可能であることが実証された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。