[논문 리뷰] Large Language Models Can Infer Psychological Dispositions of Social Media Users
본 연구는 GPT-3.5 및 GPT-4가 제로샷 설정에서 Facebook 상태 업데이트로 Big Five 성격 특성을 추론할 수 있음을 보여주며, 자기보고와의 평균 상관계수는 약 .29이고 성별 및 연령 편향이 나타난다.
Large Language Models (LLMs) demonstrate increasingly human-like abilities across a wide variety of tasks. In this paper, we investigate whether LLMs like ChatGPT can accurately infer the psychological dispositions of social media users and whether their ability to do so varies across socio-demographic groups. Specifically, we test whether GPT-3.5 and GPT-4 can derive the Big Five personality traits from users' Facebook status updates in a zero-shot learning scenario. Our results show an average correlation of r = .29 (range = [.22, .33]) between LLM-inferred and self-reported trait scores - a level of accuracy that is similar to that of supervised machine learning models specifically trained to infer personality. Our findings also highlight heterogeneity in the accuracy of personality inferences across different age groups and gender categories: predictions were found to be more accurate for women and younger individuals on several traits, suggesting a potential bias stemming from the underlying training data or differences in online self-expression. The ability of LLMs to infer psychological dispositions from user-generated text has the potential to democratize access to cheap and scalable psychometric assessments for both researchers and practitioners. On the one hand, this democratization might facilitate large-scale research of high ecological validity and spark innovation in personalized services. On the other hand, it also raises ethical concerns regarding user privacy and self-determination, highlighting the need for stringent ethical frameworks and regulation.
연구 동기 및 목표
- LLMs가 explicit training 없이 소셜 미디어 텍스트로 Big Five 성격 특성을 추론할 수 있는지 평가.
- GPT-3.5 및 GPT-4의 제로샷 추론 성능을 Facebook 상태 업데이트를 사용하여 평가.
- LLM 기반 추론에서 잠재적 인구통계학적 편향(성별 및 연령)을 고찰.
제안 방법
- 1000명의 MyPersonality 참가자와 IPIP 자기보고 및 최소 200개의 Facebook 상태 업데이트를 사용.
- 각 사용자별 마지막 200개의 상태 업데이트를 연결하고 GPT-3.5 및 GPT-4에게 열림성, 성실성, 외향성, 친화성, 신경증성 지표를 1–5 척도로 평가하도록 프롬프트한다.
- 업데이트를 20개 메시지 단위로 처리하고 평가 라운드 3회를 평균하여 전체 특성 점수를 얻는다.
- 피어슨 상관을 사용하여 LLM 추론 점수와 자기보고된 IPIP 점수를 비교한다.
- 잔차 분석을 통해 성별 및 연령 그룹 간 정확도 차이를 평가한다.
실험 결과
연구 질문
- RQ1GPT-3.5 및 GPT-4가 제로샷 설정에서 소셜 미디어 텍스트로 Big Five 성격 특성을 추론할 수 있는가?
- RQ2추론된 특성이 자기보고된 특성 점수와 어떻게 상관되며, 특성 및 모델 버전에 따라 이 상관이 어떻게 달라지는가?
- RQ3성별 또는 연령이 LLM 기반 성격 추론의 정확도나 편향에 영향을 미치는가?
- RQ4입력 텍스트의 양이 추론 정확도에 어떤 영향을 미치는가?
주요 결과
- GPT-3.5는 평균 상관 r = .27을 달성했고; GPT-4는 r = .31을 달성했다.
- 특성 수준의 상관은 Openness(.28 / .33), Extraversion(.29 / .32), Agreeableness(.30 / .32)가 각각 GPT-3.5 / GPT-4에서 가장 높았다.
- Conscientiousness는 낮은 상관(.22 / .26)과 Neuroticism은 (.26 / .29)으로 나타났다 GPT-3.5 / GPT-4에서 각각.
- 전반적으로 GPT-4가 GPT-3.5보다 더 정확한 추론을 제공했으나, 특성별 차이는 보정 후에도 통계적으로 유의미하지 않았다.
- 성별 분석은 여성의 여러 특성에서 추론된 점수가 더 높은 경향이 보이고 남성 추론의 잔차가 더 커서 남성의 여러 특성에서 정확도가 낮음을 시사한다.
- 연령 분석은 더 나이가 많은 사용자가 자기보고된 Conscientiousness 및 Neuroticism 관련 차이가 더 큰 경향을 보이며, 일부 특성에서는 모델에 따라 연령이 많을수록 정확도가 감소하는 것으로 나타났다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.