Skip to main content
QUICK REVIEW

[论文解读] Semantic features of object concepts generated with GPT-3

Hannes Hansen, Martin N. Hebart|arXiv (Cornell University)|Feb 8, 2022
Topic Modeling被引用 12
一句话总结

本研究证明,GPT-3 能够自动生成 1,854 个具体物体概念的高质量语义特征,生成的特征集在预测概念相似性、相关性和类别归属方面与人类标注的数据相当。这些生成的特征与人类知识高度一致,为认知科学研究提供了可扩展且可解释的特征规范。

ABSTRACT

Semantic features have been playing a central role in investigating the nature of our conceptual representations. Yet the enormous time and effort required to empirically sample and norm features from human raters has restricted their use to a limited set of manually curated concepts. Given recent promising developments with transformer-based language models, here we asked whether it was possible to use such models to automatically generate meaningful lists of properties for arbitrary object concepts and whether these models would produce features similar to those found in humans. To this end, we probed a GPT-3 model to generate semantic features for 1,854 objects and compared automatically-generated features to existing human feature norms. GPT-3 generated many more features than humans, yet showed a similar distribution in the types of generated features. Generated feature norms rivaled human norms in predicting similarity, relatedness, and category membership, while variance partitioning demonstrated that these predictions were driven by similar variance in humans and GPT-3. Together, these results highlight the potential of large language models to capture important facets of human knowledge and yield a new approach for automatically generating interpretable feature sets, thus drastically expanding the potential use of semantic features in psychological and linguistic studies.

研究动机与目标

  • 探究类似 GPT-3 的大语言模型是否能够自动生成任意物体概念的有意义语义特征。
  • 将 GPT-3 生成的特征质量与结构与既有的人类标注语义特征规范进行比较。
  • 评估 GPT-3 生成的特征在预测人类评分的概念相似性、相关性和类别归属方面,是否与人类规范具有同等准确性。
  • 确定 GPT-3 特征的预测能力是否源于与人类概念知识相同的潜在信息。
  • 建立一种可扩展的自动化方法,用于生成心理学和语言学研究中可解释的语义特征规范。

提出的方法

  • 采用零样本提示策略,仅使用三个 few-shot 示例,让 GPT-3 为 1,854 个具体物体概念生成语义特征。
  • 对生成的特征进行预处理和过滤,以去除重复项和低质量条目,最终得到 1,854 个概念中包含 11,683 个唯一特征的优化特征规范。
  • 通过概念间特征重叠计算语义相似性和相关性,并将预测结果与 THINGS 数据集中的人类相似性评分进行评估。
  • 应用方差划分方法,评估 GPT-3、McRae 和 CSLB 规范在解释人类相似性评分方面所共享与独特的贡献。
  • 比较 GPT-3 与人类规范之间的特征分布和标签频率,以评估结构相似性。
  • 通过现有的人类规范(McRae 和 CSLB)以及 THINGS 数据集中的相似性评分对方法进行验证,实现直接比较。

实验结果

研究问题

  • RQ1GPT-3 能否生成在定性和定量上均与人类生成规范相当的任意物体概念语义特征?
  • RQ2GPT-3 生成的特征在预测人类评分的概念相似性和相关性方面,与人类规范相比表现如何?
  • RQ3GPT-3 特征在多大程度上依赖于与人类概念知识相同的潜在信息?
  • RQ4GPT-3 生成的特征在分布和类型上与人类生成的特征相比有何异同?
  • RQ5GPT-3 是否可作为可扩展的自动化工具,用于在广泛概念范围内生成可解释的语义特征规范?

主要发现

  • GPT-3 共生成 189,126 个特征,覆盖 1,854 个概念,过滤后剩余 124,569 个,平均每个概念生成 67.35 个特征,显著高于人类规范(McRae 为 13.42,CSLB 为 35.52)。
  • GPT-3 输出中各类特征(如物理属性、功能、类别)的分布与人类规范高度一致,表明结构相似性。
  • GPT-3 生成的特征在预测 THINGS 数据集中的人类相似性评分方面,表现与 McRae 和 CSLB 人类规范相当,表现为高皮尔逊相关系数。
  • 方差划分分析显示,GPT-3 与人类规范在解释方差方面存在显著重叠,GPT-3 捕获了 McRae 规范解释的大部分方差,并与 CSLB 解释的方差部分相当。
  • 经过滤后的 GPT-3 特征(11,683 个唯一特征)在预测性能上与完整特征集几乎无差异,表明该方法具有鲁棒性和高效性。
  • 结果表明,GPT-3 抓住了人类概念知识的核心方面,尤其在支持相似性判断的共享信息方面。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。