Skip to main content
QUICK REVIEW

[论文解读] Using Large Language Models to Create AI Personas for Replication, Generalization and Prediction of Media Effects: An Empirical Test of 133 Published Experimental Research Findings

Leo Yeykelis, Kaavya Pichai|arXiv (Cornell University)|Aug 28, 2024
Persona Design and Applications被引用 4
一句话总结

本研究评估了大型语言模型(LLMs)作为AI角色,以复制、泛化并预测营销研究中的媒体效应。通过基于14本《市场营销杂志》研究中的133项实验发现生成19,447名AI参与者,LLMs成功复制了76%的主要效应(111项中的84项)和68%的所有效应(133项中的90项),展示了其在加速理论构建和提升媒体与营销心理学研究可重复性方面的巨大潜力。

ABSTRACT

This report analyzes the potential for large language models (LLMs) to expedite accurate replication and generalization of published research about message effects in marketing. LLM-powered participants (personas) were tested by replicating 133 experimental findings from 14 papers containing 45 recent studies published in the Journal of Marketing. For each study, the measures, stimuli, and sampling specifications were used to generate prompts for LLMs to act as unique personas. The AI personas, 19,447 in total across all of the studies, generated complete datasets and statistical analyses were then compared with the original human study results. The LLM replications successfully reproduced 76% of the original main effects (84 out of 111), demonstrating strong potential for AI-assisted replication. The overall replication rate including interaction effects was 68% (90 out of 133). Furthermore, a test of how human results generalized to different participant samples, media stimuli, and measures showed that replication results can change when tests go beyond the parameters of the original human studies. Implications are discussed for the replication and generalizability crises in social science, the acceleration of theory building in media and marketing psychology, and the practical advantages of rapid message testing for consumer products. Limitations of AI replications are addressed with respect to complex interaction effects, biases in AI models, and establishing benchmarks for AI metrics in marketing research.

研究动机与目标

  • 通过测试LLMs作为合成参与者的角色,解决社会科学中的可重复性与泛化危机。
  • 通过快速、可扩展的消息测试,加速媒体与营销心理学领域的理论构建。
  • 评估LLM生成数据在复制多样化刺激与测量方式下已发表实验发现的保真度。
  • 评估AI复制的局限性,特别是针对复杂交互效应和模型偏差的问题。
  • 为营销研究中的AI指标建立基准,以支持未来的方法论创新。

提出的方法

  • 从《市场营销杂志》的14篇已发表论文中提取研究参数(测量方式、刺激物、抽样规格)。
  • 通过提示工程生成独特的LLM角色,以模拟多样化的参与者特征。
  • 指导LLMs通过以真实参与者身份响应刺激物,生成完整的数据集。
  • 对AI生成的数据进行统计分析,并与原始人类研究结果进行比较。
  • 通过比较主要效应与交互效应的效应量、p值和显著性水平,评估复制成功的程度。
  • 通过在AI复制中改变参与者样本、媒体刺激物和测量工具,测试其泛化能力。

实验结果

研究问题

  • RQ1LLM生成的AI角色在多大程度上能准确复制133项已发表媒体效应实验的主要效应?
  • RQ2当参与者样本、刺激物或测量工具与原始研究不同时,LLM复制在多大程度上能实现泛化?
  • RQ3LLMs在复制媒体与营销研究中复杂交互效应方面存在哪些局限性?
  • RQ4在统计显著性与效应量方面,AI生成结果与原始人类研究结果相比如何?
  • RQ5为确保LLMs在营销研究中可靠使用,需要哪些基准与方法论标准?

主要发现

  • LLMs成功复制了所测试133项实验发现中的84项主要效应(占111项的76%)。
  • 总体复制率(包括交互效应)为68%(133项中的90项),表明保真度较强但并非完美。
  • 在改变参与者样本、媒体刺激物或测量工具以测试泛化时,复制准确性发生变化,表明AI结果会因参数变动而改变。
  • AI复制在消费者产品开发中的快速消息测试方面展现出实际优势,得益于其可扩展性与速度。
  • 在建模复杂交互效应以及LLM输出中的潜在偏差方面观察到局限性,凸显了严格验证的必要性。
  • 本研究为将LLMs作为可扩展工具用于加速理论构建并提升媒体与营销心理学研究的可重复性奠定了基础。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。