[Paper Review] Automated Social Science: Language Models as Scientist and Subjects
The paper automates generating and testing social science hypotheses in silico using LLMs guided by structural causal models (SCMs), and evaluates four scenarios (bargaining, bail, job interview, and auction).
We present an approach for automatically generating and testing, in silico, social scientific hypotheses. This automation is made possible by recent advances in large language models (LLM), but the key feature of the approach is the use of structural causal models. Structural causal models provide a language to state hypotheses, a blueprint for constructing LLM-based agents, an experimental design, and a plan for data analysis. The fitted structural causal model becomes an object available for prediction or the planning of follow-on experiments. We demonstrate the approach with several scenarios: a negotiation, a bail hearing, a job interview, and an auction. In each case, causal relationships are both proposed and tested by the system, finding evidence for some and not others. We provide evidence that the insights from these simulations of social interactions are not available to the LLM purely through direct elicitation. When given its proposed structural causal model for each scenario, the LLM is good at predicting the signs of estimated effects, but it cannot reliably predict the magnitudes of those estimates. In the auction experiment, the in silico simulation results closely match the predictions of auction theory, but elicited predictions of the clearing prices from the LLM are inaccurate. However, the LLM's predictions are dramatically improved if the model can condition on the fitted structural causal model. In short, the LLM knows more than it can (immediately) tell.
Motivation & Objective
- Formalize a workflow that uses SCMs as a blueprint to generate agents, design experiments, and analyze data with LLMs.
- Automate hypothesis generation and in silico hypothesis testing for social science questions.
- Demonstrate the approach across multiple scenarios and compare LLM predictions to theory and simulation outcomes.
Proposed method
- Represent causal relationships with simple linear SCMs to guide hypothesis generation and experimental design.
- Instantiate agents as LLM-powered entities that vary on exogenous SCM dimensions.
- Use a turn-taking protocol for agent interactions to simulate conversations and collect data.
- Run parallel simulations across exogenous dimensions and estimate the linear SCM to obtain path coefficients.
- Provide a pre-analysis plan embedded in the SCM to guide data analysis and interpretation.
- Compare LLM-predicted path signs and magnitudes to simulation estimates and theory.

Experimental results
Research questions
- RQ1Can an SCM-guided autonomous system generate and test social science hypotheses using LLM-powered agents?
- RQ2Do in silico simulations reproduce known theoretical and empirical patterns in bargaining, bail decisions, interviews, and auctions?
- RQ3Are LLMs able to predict the direction of effects and the magnitudes of effects, and how do their predictions improve when conditioned on the fitted SCM?
Key findings
- The system generated and tested falsifiable hypotheses across four scenarios, finding significant effects in several causal paths.
- In the mug bargaining scenario, buyer budget, seller minimum price, and seller love significantly affected deal probability; magnitudes were quantified.
- In the bail scenario, defendant history significantly raised bail; remorse had smaller or conditional effects.
- In the job interview scenario, passing the bar had a large positive effect on hiring; height and interviewer friendliness were not robust predictors.
- In the auction, bidder budgets positively affected final price, with magnitudes aligned with open-ascension auction theory.
- LLM-only prompts to predict path coefficients or outcomes were less accurate than the simulation results, though conditioning on the fitted SCM improved predictions.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.