Skip to main content
QUICK REVIEW

[論文レビュー] Avalon's Game of Thoughts: Battle Against Deception through Recursive Contemplation

Shenzhi Wang, Chang Liu|arXiv (Cornell University)|Oct 2, 2023
Topic Modeling被引用数 5
ひとこと要約

本稿では、Avalonなどの協力・裏切りゲームにおける誤情報の検出と対抗を支援するため、再帰的思考と多段階的視点の取り替えを模倣することで、大規模言語モデル(LLM)の性能を向上させる、新しいフレームワークであるRecursive Contemplation(ReCon)を提案する。ReConはファインチューニングを必要とせず、チェーン・オブ・チョイス推論を用いる際、善いチームの勝率を15.0%から19.4%まで向上させる。

ABSTRACT

Recent breakthroughs in large language models (LLMs) have brought remarkable success in the field of LLM-as-Agent. Nevertheless, a prevalent assumption is that the information processed by LLMs is consistently honest, neglecting the pervasive deceptive or misleading information in human society and AI-generated content. This oversight makes LLMs susceptible to malicious manipulations, potentially resulting in detrimental outcomes. This study utilizes the intricate Avalon game as a testbed to explore LLMs' potential in deceptive environments. Avalon, full of misinformation and requiring sophisticated logic, manifests as a "Game-of-Thoughts". Inspired by the efficacy of humans' recursive thinking and perspective-taking in the Avalon game, we introduce a novel framework, Recursive Contemplation (ReCon), to enhance LLMs' ability to identify and counteract deceptive information. ReCon combines formulation and refinement contemplation processes; formulation contemplation produces initial thoughts and speech, while refinement contemplation further polishes them. Additionally, we incorporate first-order and second-order perspective transitions into these processes respectively. Specifically, the first-order allows an LLM agent to infer others' mental states, and the second-order involves understanding how others perceive the agent's mental state. After integrating ReCon with different LLMs, extensive experiment results from the Avalon game indicate its efficacy in aiding LLMs to discern and maneuver around deceptive information without extra fine-tuning and data. Finally, we offer a possible explanation for the efficacy of ReCon and explore the current limitations of LLMs in terms of safety, reasoning, speaking style, and format, potentially furnishing insights for subsequent research.

研究の動機と目的

  • LLMが社会的相互作用環境における誤情報に対してどれほど脆弱であるかを調査すること。
  • 人間のような再帰的思考と視点の取り替えが、LLMの欺瞞検出能力を向上させることを検討すること。
  • 追加のファインチューニングやデータを必要とせず、LLMエージェントが欺瞞について推論できるフレームワークを開発すること。
  • 再帰的内省が、欺瞞状況における推論、安全性、倫理的整合性の向上にどの程度効果的であるかを評価すること。
  • 現在のLLMが欺瞞を扱う際の推論、話し方のスタイル、フォーマット、安全性における限界についての知見を提供すること。

提案手法

  • ReConは、2つの認知プロセスを導入する:形成的内省(初期の考えや発言の生成)と精錬的内省(正確性を高めるための考えの洗練)。
  • 一次的視点移行により、LLMは自らの視点から他者の心の状態を推測できる。
  • 二次的視点移行により、LLMは他者が自らの心の状態をどのように認識しているかをモデル化できる。
  • これらのプロセスを統合し、再帰的思考ループを構築することで、より深い認知的反省を模倣する。
  • ReConは、ファインチューニングを一切行わず、プロンプト工学と自己一貫性のある推論のみを用いてLLMに適用する。
  • 実験は、APIアクセスおよび公開済みチェックポイントを介してGPT-3.5、GPT-4、Claude-2、LLaMA-2-70b-chat-hfなどのLLMを用いて実施された。

実験結果

リサーチクエスチョン

  • RQ1再帰的思考と多段階的視点の取り替えは、協力・裏切りゲームにおけるLLMの欺瞞検出能力を向上させることができるか?
  • RQ2標準的なチェーン・オブ・チョイスプロンプトと比較して、ReConは欺瞞環境下でのLLM性能をどの程度向上させるか?
  • RQ3ファインチューニングなしで、ReConは倫理的推論をどの程度向上させ、操作への感受性を低下させることができるか?
  • RQ4欺瞞状況に適用した際、現在のLLMが推論、話し方のスタイル、フォーマット、安全性においてどのような限界を示すか?
  • RQ5ReConは、Werewolfやミステリー・アドベンチャーなどの欺瞞が豊富な環境にも一般化可能か?

主な発見

  • チェーン・オブ・チョイス推論を用いる際、ReConはAvalonゲームにおける善いチームの勝率を15.0%から19.4%まで向上させ、欺瞞の検出能力が向上していることを示している。
  • 両チームがReConを使用する場合、悪いチームの勝率は85.0%から70.6%に低下し、ReConが欺瞞に対抗する能力を向上させていることが示された。
  • ReConは、再帰的思考プロセスと多段階的視点の取り替えを可能にすることで、欺瞞環境下でのLLMの推論能力を向上させた。
  • 追加のデータやファインチューニングを必要とせず、ReConは誤魔化しやすい発言を識別し、倫理的推論に整合させる能力を向上させた。
  • 定性的分析の結果、ReConはマルチエージェントの欺瞞状況において、より一貫性があり、文脈に適した、戦略的に整合性のある推論を促進した。
  • その利点にもかかわらず、ReConはLLMの安全性、推論の一貫性、話し方のスタイルの整合性、複雑な欺瞞タスクにおけるフォーマットの耐性に限界を露呈した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。