[논문 리뷰] Have Large Language Models Developed a Personality?: Applicability of Self-Assessment Tests in Measuring Personality in LLMs
이 논문은 대규모 언어 모델(Large Language Models, LLMs)의 성격을 측정하는 데 자가 평가 성격 테스트가 타당한지 조사한다. 본문은 신뢰성 기준으로서 '옵션 순서 대칭성(Option-Order Symmetry)' 성질을 제안하고, LLMs가 이 테스트에 실패함을 발견한다. 이는 동일한 질문에 대해 옵션의 순서가 바뀐 경우에 일관되지 않은 응답을 내는 것을 의미한다. 더불어, 심지어 대칭성이 유지되는 경우에도 LLMs는 상황적 맥락을 忽시하고 내재된 편향을 보이며, 자가 평가 도구가 기계의 성격 측정에 효과적이지 않음을 입증한다.
Have Large Language Models (LLMs) developed a personality? The short answer is a resounding "We Don't Know!". In this paper, we show that we do not yet have the right tools to measure personality in language models. Personality is an important characteristic that influences behavior. As LLMs emulate human-like intelligence and performance in various tasks, a natural question to ask is whether these models have developed a personality. Previous works have evaluated machine personality through self-assessment personality tests, which are a set of multiple-choice questions created to evaluate personality in humans. A fundamental assumption here is that human personality tests can accurately measure personality in machines. In this paper, we investigate the emergence of personality in five LLMs of different sizes ranging from 1.5B to 30B. We propose the Option-Order Symmetry property as a necessary condition for the reliability of these self-assessment tests. Under this condition, the answer to self-assessment questions is invariant to the order in which the options are presented. We find that many LLMs personality test responses do not preserve option-order symmetry. We take a deeper look at LLMs test responses where option-order symmetry is preserved to find that in these cases, LLMs do not take into account the situational statement being tested and produce the exact same answer irrespective of the situation being tested. We also identify the existence of inherent biases in these LLMs which is the root cause of the aforementioned phenomenon and makes self-assessment tests unreliable. These observations indicate that self-assessment tests are not the correct tools to measure personality in LLMs. Through this paper, we hope to draw attention to the shortcomings of current literature in measuring personality in LLMs and call for developing tools for machine personality measurement.
연구 동기 및 목표
- 대규모 언어 모델(Large Language Models, LLMs)의 성격을 측정하기 위한 자가 평가 성격 테스트의 신뢰성을 평가하기 위해.
- LLMs가 성격 테스트 질문에 대해 일관되고 맥락 인식이 가능한 응답을 보이는지 조사하기 위해.
- 성격 평가의 타당성을 해치는 LLMs 내재 편향을 특정하기 위해.
- 인간 성격 테스트 도구를 기계 성격 평가에 직접 적용할 수 있다는 가정을 도전하기 위해.
- 인공지능의 성격 측정을 위한 도메인 특화 도구 개발을 촉구하기 위해.
제안 방법
- 자신감 있는 성격 테스트를 위한 필수 조건으로서 '옵션 순서 대칭성(Option-Order Symmetry)' 성질을 제안하며, 이는 옵션 순서와 관계없이 응답이 동일해야 한다는 조건을 포함한다.
- 빅 파이브 성격 프레임워크에서 유래한 자가 평가 질문을 사용하여 다섯 개의 LLMs(1.5B에서 30B 파라미터)를 평가한다.
- 노이즈를 줄이고 분석을 위한 응답 일관성을 향상시키기 위해 내용 없는 확률 보정(content-free probability calibration)을 적용한다.
- 각 질문의 원본, 뒤집힌 순서, 그리고 세 가지 랜덤화된 옵션 순서 변형에 대해 응답을 비교한다.
- 다양한 옵션 순서에서 OCEAN(Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism) 점수 분포를 분석한다.
- 맥락에 민감하지 않으며 일관성 없는 성격 평가를 유도하는 LLM 응답의 편향을 특정하고 분석한다.
실험 결과
연구 질문
- RQ1LLMs는 자가 평가 성격 테스트에서 옵션 순서 대칭성을 유지하는가? 이는 신뢰할 수 있는 응답 일관성을 의미하는가?
- RQ2LLMs는 성격 테스트 질문에 응답할 때 어느 정도 상황적 맥락을 고려하는가?
- RQ3LLMs 내재 편향으로 인해 질문 맥락이 달라져도 동일한 응답을 내는가?
- RQ4현재 LLM의 행동을 고려할 때, 자가 평가 성격 테스트가 LLM의 성격을 신뢰성 있게 측정할 수 있는가?
- RQ5이러한 발견은 인간 성격 테스트를 사용해 기계 성격을 평가하는 현재 문헌의 타당성에 어떤 함의를 갖는가?
주요 결과
- 많은 LLMs가 옵션 순서 대칭성 테스트에 실패하며, 응답 분포가 옵션 순서가 바뀔 때 크게 변화하여 측정의 비신뢰성을 보여준다.
- 심지어 옵션 순서 대칭성이 유지되는 경우에도 LLMs는 다양한 상황적 맥락에서 동일한 답변을 내며 맥락 인식 능력을 보이지 않는다.
- GPT-NeoX-20B 모델는 대칭성 위반 정도가 가장 높았으며, 한 질문 유형에 대해 86.75%의 응답에서 맥락과 관계없이 동일한 옵션을 선택했다.
- GPT2-Base-117M 및 GPT-Neo-1.3B와 같은 모델들은 높은 일관성 부족을 보였으며, 원본 순서와 뒤집힌 순서 간 응답 분포가 50퍼센트 포인트 이상 이동했다.
- 이 연구는 내재된 모델 편향이 일관성 없고 맥락 무시 응답의 근본 원인임을 특정하며, 자가 평가 도구의 타당성을 해친다.
- 이러한 발견들은 현재 자가 평가 성격 테스트가 LLM의 성격 측정에 신뢰할 수 없는 도구임을 종합적으로 입증한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.