Skip to main content
QUICK REVIEW

[论文解读] Measurement in the Age of LLMs: An Application to Ideological Scaling

Sean O’Hagan, Aaron Schein|arXiv (Cornell University)|Dec 14, 2023
Computational and Text Analysis Methods被引用 5
一句话总结

本文提出使用大语言模型(LLMs)直接提取文本和个体的数值意识形态得分,绕过传统数据限制。通过提示GPT-3.5-turbo采用零样本思维链推理方式分配得分,该方法与既有的意识形态量化方法具有高度相关性(r = 0.94),表明LLMs能够捕捉自然语言中细微且弥散的意识形态信号。

ABSTRACT

Much of social science is centered around terms like ``ideology'' or ``power'', which generally elude precise definition, and whose contextual meanings are trapped in surrounding language. This paper explores the use of large language models (LLMs) to flexibly navigate the conceptual clutter inherent to social scientific measurement tasks. We rely on LLMs' remarkable linguistic fluency to elicit ideological scales of both legislators and text, which accord closely to established methods and our own judgement. A key aspect of our approach is that we elicit such scores directly, instructing the LLM to furnish numeric scores itself. This approach affords a great deal of flexibility, which we showcase through a variety of different case studies. Our results suggest that LLMs can be used to characterize highly subtle and diffuse manifestations of political ideology in text.

研究动机与目标

  • 解决测量抽象社会科学概念(如'意识形态')的挑战,这些概念缺乏精确定义且嵌入于语言语境之中。
  • 探索大语言模型(LLMs)是否可作为协作性、灵活的工具用于社会科学测量,而无需预先定义操作化标准。
  • 评估直接提示LLM(尤其是零样本思维链)在获取一致、可解释的意识形态得分方面的有效性。
  • 在相关性和一致性方面,将基于LLM的意识形态量化方法与既有的DW-NOMINATE、TBIP和CFScores方法进行比较。
  • 评估在LLM驱动的测量任务中,初始的二元判断是否改善或损害性能。

提出的方法

  • 使用GPT-3.5-turbo直接为文本或个体分配数值意识形态得分(连续尺度),将LLM视为协作对话对象。
  • 通过人类反馈实现动态自我锚定,以在不同输入类型下稳定并校准LLM的响应。
  • 应用零样本思维链提示,引导LLM在分配最终数值得分前进行逐步推理。
  • 测试三种提示策略:(1) 直接对所有推文打分,(2) 仅对经二元判断为具有意识形态内容的推文打分,(3) 结合二元判断与零样本思维链。
  • 使用核密度估计和相关性分析,将LLM获取的得分与DW-NOMINATE、TBIP和CFScores的参考意识形态量表进行比较。
  • 整体评估性能,并在政党群体(共和党和民主党)内评估,以衡量对语言风格和意识形态细微差别的敏感性。
Figure 1 : Kernel density estimates of the distribution of GPT-elicited ideal points for Donald Trump and Bernie Sanders. Two prompts are used: one which leaves open the interpretation of “authoritarian-libertarian”, and one which elaborates informally.
Figure 1 : Kernel density estimates of the distribution of GPT-elicited ideal points for Donald Trump and Bernie Sanders. Two prompts are used: one which leaves open the interpretation of “authoritarian-libertarian”, and one which elaborates informally.

实验结果

研究问题

  • RQ1LLMs能否在不依赖结构化行为数据的情况下,可靠地为文本和个体提取数值意识形态得分?
  • RQ2引入零样本思维链推理如何影响基于LLM的意识形态量化的稳定性和准确性?
  • RQ3通过二元判断对意识形态内容进行预筛选,是否能提升或阻碍基于LLM的测量性能?
  • RQ4LLM获取的意识形态得分与既有的参考方法(如DW-NOMINATE和TBIP)之间的相关性如何?
  • RQ5LLMs在多大程度上能够捕捉政治话语中细微且弥散的意识形态信号,特别是在共和党推文等党派语境中?

主要发现

  • 基于LLM的方法在最终打分方法(结合零样本思维链与意识形态过滤)与GPT获取的理想点之间实现了0.94的相关性,表明其具有极强的内部一致性。
  • 在共和党群体中,若不使用零样本思维链而仅通过二元过滤,LLM得分与参考方法的相关性显著下降(r = 0.0),表明此类过滤可能损害对细微、弥散意识形态内容的性能。
  • 采用零样本思维链与意识形态过滤的方法(图7中的深蓝色)在所考察的五位参议员中,均表现出最集中且最准确的得分分布,围绕GPT获取的理想点展开。
  • 核密度估计显示,仅最复杂的提示策略(结合零样本思维链)产生了与GPT获取的理想点一致的得分分布,尤其在推文意识形态复杂度较高的参议员中表现突出。
  • 结果表明,LLMs能够有效测量文本中细微的意识形态细微差别,尤其在基于推理的提示引导下,即使在意识形态未明确表达的语境中亦然。
  • 本研究证明,LLMs可作为可扩展、灵活的社会科学测量工具,保留诸如意识形态等概念的语言丰富性,而无需依赖结构化行为数据源。
Figure 2 : Ideal point scores for members of the $114^{\textrm{th}}$ U.S. Senate. The GPT-elicited scores are depicted as box-plots, which describe variability across variations in prompt (i.e., Senator order). Rows are sorted in in ascending order by the mean of the GPP-elicited scores. We compare
Figure 2 : Ideal point scores for members of the $114^{\textrm{th}}$ U.S. Senate. The GPT-elicited scores are depicted as box-plots, which describe variability across variations in prompt (i.e., Senator order). Rows are sorted in in ascending order by the mean of the GPP-elicited scores. We compare

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。