Skip to main content
QUICK REVIEW

[论文解读] Large Language Models Can Infer Psychological Dispositions of Social Media Users

Heinrich Peters, Sandra Matz|arXiv (Cornell University)|Sep 13, 2023
Artificial Intelligence in Healthcare and Education被引用 13
一句话总结

该研究表明 GPT-3.5 和 GPT-4 能在零-shot 设置从 Facebook 状态更新推断大五人格,在自我报告的平均相关性约为 .29,并且表现出性别和年龄偏差。

ABSTRACT

Large Language Models (LLMs) demonstrate increasingly human-like abilities across a wide variety of tasks. In this paper, we investigate whether LLMs like ChatGPT can accurately infer the psychological dispositions of social media users and whether their ability to do so varies across socio-demographic groups. Specifically, we test whether GPT-3.5 and GPT-4 can derive the Big Five personality traits from users' Facebook status updates in a zero-shot learning scenario. Our results show an average correlation of r = .29 (range = [.22, .33]) between LLM-inferred and self-reported trait scores - a level of accuracy that is similar to that of supervised machine learning models specifically trained to infer personality. Our findings also highlight heterogeneity in the accuracy of personality inferences across different age groups and gender categories: predictions were found to be more accurate for women and younger individuals on several traits, suggesting a potential bias stemming from the underlying training data or differences in online self-expression. The ability of LLMs to infer psychological dispositions from user-generated text has the potential to democratize access to cheap and scalable psychometric assessments for both researchers and practitioners. On the one hand, this democratization might facilitate large-scale research of high ecological validity and spark innovation in personalized services. On the other hand, it also raises ethical concerns regarding user privacy and self-determination, highlighting the need for stringent ethical frameworks and regulation.

研究动机与目标

  • 评估在没有显式训练的情况下,LLMs 是否能够从社交媒体文本推断大五人格特质。
  • 使用 Facebook 状态更新评估 GPT-3.5 和 GPT-4 的零-shot 推断性能。
  • 检查 LLM 基于推断中的潜在人口统计偏差(性别和年龄)。

提出的方法

  • 使用 1000 名 MyPersonality 参与者,具备 IPIP 自我报告并且至少有 200 条 Facebook 状态更新。
  • 将每个用户的最近 200 条状态更新拼接起来,提示 GPT-3.5 和 GPT-4 对 开放性、尽责性、外向性、宜人性、神经质性按 1–5 的量表进行评分。
  • 将更新分成 20 条消息的区块进行处理,并对三轮评分取平均以获得汇总的人格特质分数。
  • 用 Pearson 相关性将 LLM 推断得分与自报 IPIP 得分进行比较。
  • 通过残差分析评估不同性别和年龄组的准确性差异。

实验结果

研究问题

  • RQ1GPT-3.5 和 GPT-4 能否在零-shot 设置从社交媒体文本推断大五人格特质?
  • RQ2推断的特质与自报特质分数之间的相关性如何?并且不同特质和模型版本会有何差异?
  • RQ3性别或年龄是否会影响基于 LLM 的人格推断的准确性或偏差?
  • RQ4输入文本量如何影响推断准确性?

主要发现

  • GPT-3.5 实现了平均相关 r = .27;GPT-4 实现了 r = .31,覆盖所有特质。
  • 在特质层面,开朗性(Openness)为最高相关性(.28 / .33),外向性(Extraversion)为最高相关性(.29 / .32),宜人性(Agreeableness)为最高相关性(.30 / .32),分别对应 GPT-3.5 / GPT-4。
  • 慎重性(Conscientiousness)的相关性较低(.22 / .26),神经质性(Neuroticism)为(.26 / .29),分别对应 GPT-3.5 / GPT-4。
  • 总体而言,GPT-4 提供的推断比 GPT-3.5 更准确,尽管在修正后没有任何特质差异达到统计显著性。
  • 性别分析显示女性在若干特质上表现出较高的推断分数,而男性推断的残差较大,表明在多项特质上的男性准确性较低。
  • 年龄分析表明年龄较大的用户自报的认真性和与神经质相关的差异较大,某些特质在较老用户上在不同模型中表现出较低的准确性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。