[论文解读] Automated Social Science: Language Models as Scientist and Subjects
本论文利用结构因果模型(SCMs)引导的LLMs,在计算机内自动生成并测试社会科学假设,并评估四种情景(谈判、保释、工作面试和拍卖)。
We present an approach for automatically generating and testing, in silico, social scientific hypotheses. This automation is made possible by recent advances in large language models (LLM), but the key feature of the approach is the use of structural causal models. Structural causal models provide a language to state hypotheses, a blueprint for constructing LLM-based agents, an experimental design, and a plan for data analysis. The fitted structural causal model becomes an object available for prediction or the planning of follow-on experiments. We demonstrate the approach with several scenarios: a negotiation, a bail hearing, a job interview, and an auction. In each case, causal relationships are both proposed and tested by the system, finding evidence for some and not others. We provide evidence that the insights from these simulations of social interactions are not available to the LLM purely through direct elicitation. When given its proposed structural causal model for each scenario, the LLM is good at predicting the signs of estimated effects, but it cannot reliably predict the magnitudes of those estimates. In the auction experiment, the in silico simulation results closely match the predictions of auction theory, but elicited predictions of the clearing prices from the LLM are inaccurate. However, the LLM's predictions are dramatically improved if the model can condition on the fitted structural causal model. In short, the LLM knows more than it can (immediately) tell.
研究动机与目标
- 将以 SCMs 作为蓝图来生成代理、设计实验并使用 LLMs 分析数据的工作流形式化。
- 自动化生成假设并在计算机内对社会科学问题进行假设检验。
- 在多种情景中演示该方法,并将 LLM 的预测与理论和仿真结果进行比较。
提出的方法
- 用简单的线性 SCM 来表示因果关系,以引导假设生成和实验设计。
- 将代理实体实例化为以 LLM 为驱动的实体,其在外生 SCM 维度上变化。
- 使用轮流交互协议来模拟对话并收集数据。
- 在外生维度上运行并行仿真,并估计线性 SCM 以获取路径系数。
- 在 SCM 中嵌入预分析计划,以指导数据分析和解释。
- 将 LLM 预测的路径符号和大小与仿真估计和理论进行比较。

实验结果
研究问题
- RQ1在 SCM 指导的自主系统中,能否使用以 LLM 驱动的代理生成并测试社会科学假设?
- RQ2在计算机内的仿真是否能再现谈判、保释决策、面试和拍卖等领域的已知理论与经验模式?
- RQ3LLMs 能否预测效应的方向和大小?当以拟合的 SCM 条件化时,它们的预测如何改善?
主要发现
- 该系统在四个情景中生成并测试可证伪的假设,在若干因果路径上发现显著效应。
- 在 mug 谈判情景中,买家预算、卖家最低价和卖家偏好显著影响交易概率;大小量化。
- 在保释情景中,被告历史显著提高保释金额;悔意影响较小或有条件。
- 在工作面试情景中,通过律师资格考试对雇佣有显著正向影响;身高和面试官友好度不是稳健的预测因子。
- 在拍卖中,买家预算对最终价格有正向影响,且幅度与开放升序拍卖理论一致。
- 仅用 LLM 的提示来预测路径系数或结果的准确性不如仿真结果,尽管在拟合的 SCM 条件化下预测有所改进。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。