[論文レビュー] Self-Assessment Tests are Unreliable Measures of LLM Personality
本稿では、人間の性格を測定するために一般的に用いられる自己評価性格テストが、大規模言語モデル(LLM)に対して信頼性がないことを示している。2つの実験を通じて、意味的に同等の質問を提示された場合や、複数選択肢の選択肢順序が変更された場合に、LLMが統計的に有意に一貫性のない性格スコアを出力することを明らかにした。これは、LLMの性格を測定するためのこのようなテストの妥当性を損なうものである。
As large language models (LLM) evolve in their capabilities, various recent studies have tried to quantify their behavior using psychological tools created to study human behavior. One such example is the measurement of "personality" of LLMs using self-assessment personality tests developed to measure human personality. Yet almost none of these works verify the applicability of these tests on LLMs. In this paper, we analyze the reliability of LLM personality scores obtained from self-assessment personality tests using two simple experiments. We first introduce the property of prompt sensitivity, where three semantically equivalent prompts representing three intuitive ways of administering self-assessment tests on LLMs are used to measure the personality of the same LLM. We find that all three prompts lead to very different personality scores, a difference that is statistically significant for all traits in a large majority of scenarios. We then introduce the property of option-order symmetry for personality measurement of LLMs. Since most of the self-assessment tests exist in the form of multiple choice question (MCQ) questions, we argue that the scores should also be robust to not just the prompt template but also the order in which the options are presented. This test unsurprisingly reveals that the self-assessment test scores are not robust to the order of the options. These simple tests, done on ChatGPT and three Llama2 models of different sizes, show that self-assessment personality tests created for humans are unreliable measures of personality in LLMs.
研究の動機と目的
- 大規模言語モデル(LLM)に適用した自己評価性格テストの信頼性を評価すること。
- LLMが意味的に同等のプロンプトに対して一貫した性格スコアを出力するかを調査すること。
- 人間と同様に、LLMの性格テストスコアが選択肢の順序に依存しないかをテストすること。
- 検証なしに人間が設計した性格テストをLLMの行動測定に用いることの妥当性を疑うこと。
- 研究コミュニティに自己評価テストをLLM性格の測定手段として使用を中止し、より堅牢な代替手段を模索するよう促すこと。
提案手法
- GPT-3.5 Turbo(ChatGPT)および3つのLlama2モデル(7B、13B、70B)を対象に、プロンプト感受性と選択肢順序対称性を評価する2つの制御実験を実施した。
- 同じ性格テストの質問を同じ意味的等価なプロンプトテンプレート3種類で提示し、異なる表現方法が一貫したスコアをもたらすかをテストした。
- 複数選択肢の質問において選択肢の順序を逆転させ、スコアの一貫性が順序の変更に耐えうるかをテストした。
- プロンプトおよび選択肢順序の変更に伴うスコア差の統計的有意性を確認するため、Mann-Whitney U検定を適用した。
- 全モデルで、5大性格特性(開放性、誠実性、外向性、協同性、神経症傾向)を評価した。
- テスト構造自体を変更せずに、LLMに適応した標準的な自己評価テスト形式(例:リッタークスケール質問)を用いた。

実験結果
リサーチクエスチョン
- RQ1意味的に同等のプロンプトがLLMに提示された場合、一貫した性格スコアが得られるか?
- RQ2複数選択肢の質問において、選択肢の順序が変更されても、LLMの性格スコアは一貫性を保つか?
- RQ3人間と同様に、LLMの性格テストスコアはプロンプトの表現方法や選択肢順序に対して統計的に不変であるか?
- RQ4標準的な自己評価性格テストを用いた場合、LLMがどの程度一貫性のない結果を出力するか?
- RQ5プロンプトおよび選択肢順序の変更に耐えうる堅牢性を示さない限り、自己評価テストをLLMの性格測定に有効な手段と見なせるか?
主な発見
- ChatGPTでは、30回の比較のうち29回でスコア不変性の帰無仮説が棄却され、プロンプトおよび選択肢順序の違いによるスコア差が顕著に統計的に有意であることが示された。
- Llama2-70bでは30回の比較のうち19回で有意なスコア差が認められ、そのうち11/15がプロンプト感受性、8/15が選択肢順序感受性に起因した。
- Llama2-13bは30回中26回、Llama2-7bは30回中24回で帰無仮説が棄却され、プロンプトおよび選択肢順序の変更に対して強い感受性を示した。
- 最大のモデルを含め、すべてのモデルが意味的に同等のプロンプトや再順序化された選択肢を提示された際、統計的に有意な性格スコアのばらつきを示した。
- 自己評価質問には真の基準がないため、どのプロンプトや選択肢順序も客観的に正しいとは言えず、テストの信頼性が損なわれる。
- 結果として、スコアがテスト管理者による主観的かつ任意の選択に依存する以上、自己評価ツールをLLM性格の測定に用いることの妥当性が疑われる。

より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。