[Paper Review] Using Large Language Models to Create AI Personas for Replication, Generalization and Prediction of Media Effects: An Empirical Test of 133 Published Experimental Research Findings
This study evaluates large language models (LLMs) as AI personas to replicate, generalize, and predict media effects in marketing research. By generating 19,447 AI participants based on 133 experimental findings from 14 Journal of Marketing studies, LLMs successfully replicated 76% of main effects (84/111) and 68% of all effects (90/133), demonstrating strong potential for accelerating theory building and improving replication in media and marketing psychology.
This report analyzes the potential for large language models (LLMs) to expedite accurate replication and generalization of published research about message effects in marketing. LLM-powered participants (personas) were tested by replicating 133 experimental findings from 14 papers containing 45 recent studies published in the Journal of Marketing. For each study, the measures, stimuli, and sampling specifications were used to generate prompts for LLMs to act as unique personas. The AI personas, 19,447 in total across all of the studies, generated complete datasets and statistical analyses were then compared with the original human study results. The LLM replications successfully reproduced 76% of the original main effects (84 out of 111), demonstrating strong potential for AI-assisted replication. The overall replication rate including interaction effects was 68% (90 out of 133). Furthermore, a test of how human results generalized to different participant samples, media stimuli, and measures showed that replication results can change when tests go beyond the parameters of the original human studies. Implications are discussed for the replication and generalizability crises in social science, the acceleration of theory building in media and marketing psychology, and the practical advantages of rapid message testing for consumer products. Limitations of AI replications are addressed with respect to complex interaction effects, biases in AI models, and establishing benchmarks for AI metrics in marketing research.
Motivation & Objective
- Address the replication and generalizability crisis in social science by testing LLMs as synthetic participants.
- Accelerate theory building in media and marketing psychology through rapid, scalable message testing.
- Evaluate the fidelity of LLM-generated data in reproducing published experimental findings across diverse stimuli and measures.
- Assess limitations of AI replications, particularly regarding complex interaction effects and model biases.
- Establish benchmarks for AI metrics in marketing research to support future methodological innovation.
Proposed method
- Extracted study parameters (measures, stimuli, sampling specs) from 14 published papers in the Journal of Marketing.
- Generated unique LLM personas using prompt engineering to simulate diverse participant profiles.
- Directed LLMs to produce complete datasets by responding to stimuli as if they were real participants.
- Conducted statistical analyses on AI-generated data and compared results with original human study outcomes.
- Evaluated replication success by comparing effect sizes, p-values, and significance levels across main and interaction effects.
- Tested generalization by varying participant samples, media stimuli, and measurement tools in AI replications.
Experimental results
Research questions
- RQ1To what extent can LLM-generated AI personas accurately replicate the main effects of 133 published media effects experiments?
- RQ2How well do LLM replications generalize when participant samples, stimuli, or measurement tools are altered from original studies?
- RQ3What are the limitations of LLMs in replicating complex interaction effects in media and marketing research?
- RQ4How do AI-generated results compare to original human study results in terms of statistical significance and effect size?
- RQ5What benchmarks and methodological standards are needed to ensure reliable use of LLMs in marketing research?
Key findings
- LLMs successfully replicated 84 out of 111 main effects (76%) from the 133 experimental findings tested.
- The overall replication rate, including interaction effects, was 68% (90 out of 133 effects), indicating strong but not perfect fidelity.
- Replication accuracy varied when testing generalization across different participant samples, media stimuli, or measurement tools, showing that AI results can change under parameter shifts.
- AI replications demonstrated practical advantages for rapid message testing in consumer product development due to scalability and speed.
- Limitations were observed in modeling complex interaction effects and potential biases in LLM outputs, highlighting the need for rigorous validation.
- The study establishes a foundation for using LLMs as a scalable tool to accelerate theory building and improve reproducibility in media and marketing psychology.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.