Skip to main content
QUICK REVIEW

[論文レビュー] Instruction-tuning Aligns LLMs to the Human Brain

Khai Loong Aw, Syrielle Montariol|arXiv (Cornell University)|Dec 1, 2023
Topic Modeling被引用数 5
ひとこと要約

この論文は、指令微調整(instruction-tuning)が大規模言語モデル(LLMs)と人間の脳の間の整合性を向上させるかを調査している。fMRI研究からの脳活動データと、読解時間からの行動データを用いて、指令微調整を施したLLMsは、内部表現が人間の神経応答をどれだけうまく予測できるかという指標で、平均6.2%の向上を示した。主な要因は世界知識の向上とモデルサイズの増大であり、行動的整合性はほとんど変化しなかった。

ABSTRACT

Instruction-tuning is a widely adopted finetuning method that enables large language models (LLMs) to generate output that more closely resembles human responses. However, no studies have shown that instruction-tuning actually teaches LLMs to process language in a similar manner as humans. We investigate the effect of instruction-tuning on aligning LLM and human language processing mechanisms in two ways: (1) brain alignment, the similarity of LLM internal representations to neural activity in the human language system, and (2) behavioral alignment, the similarity of LLM and human behavior on a reading task. We assess 25 vanilla and instruction-tuned LLMs on three datasets involving humans reading naturalistic stories and sentences, and find that instruction-tuning generally enhances brain alignment (~6%), but has no similar effect on behavioral alignment. To identify factors underlying this improvement in brain alignment, we compute correlations between brain alignment and various LLM properties, such as model size, problem-solving, and world knowledge understanding. Notably, we find a strong positive correlation between brain alignment and model size (r = 0.95), as well as performance on tasks requiring world knowledge (r = 0.81). Our results demonstrate that instruction-tuning LLMs improves both world knowledge representations and brain alignment, suggesting that the mechanisms that encode world knowledge in LLMs also improve representational alignment to the human brain.

研究の動機と目的

  • 指令微調整が、神経表現と行動の両面でLLMsを人間の言語処理に近づけるかどうかを検証すること。
  • 自然な言語理解の過程で、LLMsと人間のfMRI反応との脳整合性を評価すること。
  • 人間参加者からの単語ごとの読解時間データを用いて、行動的整合性を評価すること。
  • LLM-脳整合性を最も強く予測するモデル特性を同定すること。
  • 指令微調整が、出力の模倣を超えて、LLMsの代表的類似性を人間の脳に高めるかどうかを調査すること。

提案手法

  • Pereira et al. (2018)、Blank et al. (2014)、Wehbe et al. (2014) の3つのfMRIデータセットを用い、25のLLMs(8つはヴァナイル、17つは指令微調整済み)を評価した。
  • Brain-Scoreという線形予測性指標を用いて、LLMの表現が人間のfMRI活動をどれだけうまく予測できるか、脳整合性を測定した。
  • Futrell et al. (2018) データセットからのLLMの単語ごとの周辺度と人間の単語ごとの読解時間のピアソン相関を用いて、行動的整合性を評価した。
  • 次単語予測(NWP)性能、モデルサイズ、BBH(Big-Bench Hard)およびMMLU(Massive Multi-task Language Understanding)ベンチマークでのパフォーマンスを含む、モデル特性を計算した。
  • LLM特性と脳整合性の関係を分析するためにピアソン相関を用いた。
  • 完全な再現可能性を確保するため、オープンソースモデルと公開リポジトリ(Brain-Scoreおよびinstruct-eval)を用いて結果を再現した。

実験結果

リサーチクエスチョン

  • RQ1指令微調整は、言語処理中のLLM内部表現と人間の脳活動との類似性を向上させるか?
  • RQ2指令微調整は、読解課題におけるLLMsと人間の行動的類似性を向上させるか?
  • RQ3モデルサイズ、世界知識、推論能力といったモデル特性のうち、どれがLLM-脳整合性を最も強く予測するか?
  • RQ4脳整合性の向上は、世界知識の向上によるものか、それとも他の能力によるものか?
  • RQ5LLMsの推論および知識ベンチマークでのパフォーマンスは、人間の神経応答への整合性とどの程度相関するか?

主な発見

  • 指令微調整は、テストされたLLMsおよび神経データセット全体で、平均して脳整合性を6.2%向上させた。
  • モデルサイズと脳整合性の間に強い正の相関関係が見られた(r = 0.95)。
  • MMLUパフォーマンスで測定した世界知識は、脳整合性と高い正の相関を示した(r = 0.81)。
  • 指令微調整は、脳整合性に影響を与える一方で、人間の読解時間との行動的整合性に顕著な改善をもたらさなかった。
  • 現代のTransformerベースのLLMsでは、次単語予測能力と脳整合性の間には一貫した相関関係が見られなかった。
  • 結果から、世界知識とモデルスケーリングが、人間の脳への代表的類似性の主な駆動要因であると考えられる。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。