[논문 리뷰] Boosting Theory-of-Mind Performance in Large Language Models via Prompting
본 논문은 프롬프팅이 특히 두-shot chain-of-thought 또는 step-by-step 프롬프트를 포함한 인-context 학습일 때 RLHF-학습된 LLM의 이론 마음(ToM) 성능을 향상시킨다고 보여주며, GPT-4는 프롬프트에서 ToM 정확도 100%에 도달하고 제로샷 GPT-4는 약 80%에 근접한다; 인간 정확도는 87%이다.
Large language models (LLMs) excel in many tasks in 2023, but they still face challenges in complex reasoning. Theory-of-mind (ToM) tasks, which require understanding agents' beliefs, goals, and mental states, are essential for common-sense reasoning involving humans, making it crucial to enhance LLM performance in this area. This study measures the ToM performance of GPT-4 and three GPT-3.5 variants (Davinci-2, Davinci-3, GPT-3.5-Turbo), and investigates the effectiveness of in-context learning in improving their ToM comprehension. We evaluated prompts featuring two-shot chain of thought reasoning and step-by-step thinking instructions. We found that LLMs trained with Reinforcement Learning from Human Feedback (RLHF) (all models excluding Davinci-2) improved their ToM accuracy via in-context learning. GPT-4 performed best in zero-shot settings, reaching nearly 80% ToM accuracy, but still fell short of the 87% human accuracy on the test set. However, when supplied with prompts for in-context learning, all RLHF-trained LLMs exceeded 80% ToM accuracy, with GPT-4 reaching 100%. These results demonstrate that appropriate prompting enhances LLM ToM reasoning, and they underscore the context-dependent nature of LLM cognitive capacities.
연구 동기 및 목표
- Assess ToM performance of GPT-4 and three GPT-3.5 variants on ToM tasks.
- Evaluate the effect of in-context learning prompts on ToM accuracy.
- Examine differences between RLHF-trained models and non-RLHF baselines on ToM tasks.
제안 방법
- Evaluate ToM performance of GPT-4 and three GPT-3.5 variants (Davinci-2, Davinci-3, GPT-3.5-Turbo).
- Test zero-shot and in-context learning prompts including two-shot chain-of-thought and step-by-step thinking instructions.
- Compare RLHF-trained models against non-RLHF baselines regarding ToM accuracy.
- Measure ToM accuracy against human performance as a benchmark.
실험 결과
연구 질문
- RQ1How does prompting influence theory-of-mind accuracy in large language models?
- RQ2Do RLHF-trained LLMs benefit from in-context learning for ToM tasks more than non-RLHF models?
- RQ3What are the best prompting configurations (zero-shot vs. in-context, chain-of-thought vs. step-by-step) for ToM in LLMs?
- RQ4How close do LLM ToM performances come to human accuracy on the test set?
주요 결과
- GPT-4 achieves nearly 80% ToM accuracy in zero-shot settings.
- RLHF-trained models (excluding Davinci-2) improve ToM accuracy via in-context learning.
- With prompting, all RLHF-trained LLMs exceed 80% ToM accuracy; GPT-4 with prompts reaches 100%.
- GPT-4 in zero-shot approaches approaches but does not reach the human accuracy of 87% on the test set.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.