Skip to main content
QUICK REVIEW

[论文解读] Can ChatGPT Read Who You Are?

Erik Derner, Dalibor Kučera|arXiv (Cornell University)|Dec 26, 2023
Mental Health via Writing参考文献 38被引用 4
一句话总结

本研究调查了ChatGPT是否能够基于捷克语短文本,利用大五人格量表(BFI)作为真实值,推断人格特质。研究结果表明,ChatGPT在与人类评分者对比时表现具有竞争力,但其在所有特质上均表现出一致的积极偏见,且性能显著受提示设计影响。

ABSTRACT

The interplay between artificial intelligence (AI) and psychology, particularly in personality assessment, represents an important emerging area of research. Accurate personality trait estimation is crucial not only for enhancing personalization in human-computer interaction but also for a wide variety of applications ranging from mental health to education. This paper analyzes the capability of a generic chatbot, ChatGPT, to effectively infer personality traits from short texts. We report the results of a comprehensive user study featuring texts written in Czech by a representative population sample of 155 participants. Their self-assessments based on the Big Five Inventory (BFI) questionnaire serve as the ground truth. We compare the personality trait estimations made by ChatGPT against those by human raters and report ChatGPT's competitive performance in inferring personality traits from text. We also uncover a 'positivity bias' in ChatGPT's assessments across all personality dimensions and explore the impact of prompt composition on accuracy. This work contributes to the understanding of AI capabilities in psychological assessment, highlighting both the potential and limitations of using large language models for personality inference. Our research underscores the importance of responsible AI development, considering ethical implications such as privacy, consent, autonomy, and bias in AI applications.

研究动机与目标

  • 评估ChatGPT在非英语语言(捷克语)的短文本中推断人格特质的能力。
  • 以自报的大五人格量表(BFI)评估作为真实值,将ChatGPT的人格特质估计结果与人类评分者进行比较。
  • 调查大语言模型(LLM)人格评估中是否存在积极偏见及其影响。
  • 分析不同提示设计对ChatGPT人格推断准确度的影响。
  • 强调在AI驱动的用户建模中,隐私、知情同意、自主权和偏见等伦理影响。

提出的方法

  • 开展了一项用户研究,招募155名捷克语母语参与者,撰写四种类型的短文本,并完成大五人格量表(BFI)以进行人格评估。
  • 收集参与者的自评数据作为人格特质在五个维度(外向性、宜人性、尽责性、神经质性和开放性)上的真实值参考。
  • 采用多种提示变体评估ChatGPT的表现,包括零样本、 few-shot 和思维链(chain-of-thought)提示策略。
  • 使用描述性统计量(均值、中位数、众数、标准差、变异系数)评估并比较ChatGPT预测结果与人类评分者及自评数据的分布。
  • 分析ChatGPT在每个人格维度的三个水平(低、中性、高)上的预测一致性与准确性。
  • 将ChatGPT的输出与人类评分者(H_A 和 H_B)及其他大语言模型变体(如 GPT-TL、GPT-DTL 等)的输出进行比较,以评估相对表现与偏见模式。
Figure 1: Relative frequencies for all assessment scores in the low, neutral, and high spectra for all personality dimensions.
Figure 1: Relative frequencies for all assessment scores in the low, neutral, and high spectra for all personality dimensions.

实验结果

研究问题

  • RQ1ChatGPT能否基于自报的BFI评估结果,准确推断出捷克语短文本中的人格特质?
  • RQ2ChatGPT在人格推断方面与人类评分者的表现相比如何?
  • RQ3ChatGPT在其所有五个维度的人格特质估计中是否表现出系统性的积极偏见?
  • RQ4提示设计在多大程度上影响了ChatGPT人格推断的准确度?
  • RQ5在实际应用中,使用类似ChatGPT的大语言模型进行自动人格评估具有哪些伦理影响?

主要发现

  • ChatGPT在从短捷克语文本推断人格特质方面表现出具有竞争力的性能,多数维度的均值预测结果与自评数据高度一致。
  • ChatGPT的评估中观察到一致的积极偏见:在所有五个个性特质中,预测得分系统性地高于自评得分,尤其在低分和中性范围更为明显。
  • 在自评高分组中,神经质的平均预测得分为1.805,表明其对高神经质倾向存在明显低估;而开放性的平均预测得分为2.041,提示其对高开放性也存在低估。
  • 提示设计显著影响性能:few-shot 和思维链提示变体(如 GPT-DLT)相比零样本基线表现更优,其中 GPT-DLT 在高自评组中开放性的平均得分为2.789。
  • 在大多数维度上,人类评分者(H_A 和 H_B)与自评数据的一致性高于ChatGPT,尤其在神经质和开放性维度上,ChatGPT的平均预测值持续低于自评报告。
  • ChatGPT预测结果的变异系数通常低于人类评分者,表明其响应更一致但变异性更低,这可能反映出对个体差异的细微敏感性不足。
Figure 2: F1 scores for all evaluation methods and all personality dimensions (higher is better). The self-assessment score ( A ${}_{\textbf{S}}$ ) serves as the ground truth.
Figure 2: F1 scores for all evaluation methods and all personality dimensions (higher is better). The self-assessment score ( A ${}_{\textbf{S}}$ ) serves as the ground truth.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。