[논문 리뷰] Automated Social Science: Language Models as Scientist and Subjects
이 논문은 구조적 인과 모델(SCMs)에 의해 안내되는 LLMs를 사용하여 사회과학 가설을 인실리코(in silico)로 자동으로 생성하고 테스트하며, 네 가지 시나리오(협상, 보석금 결정, 면접, 경매)를 평가한다.
We present an approach for automatically generating and testing, in silico, social scientific hypotheses. This automation is made possible by recent advances in large language models (LLM), but the key feature of the approach is the use of structural causal models. Structural causal models provide a language to state hypotheses, a blueprint for constructing LLM-based agents, an experimental design, and a plan for data analysis. The fitted structural causal model becomes an object available for prediction or the planning of follow-on experiments. We demonstrate the approach with several scenarios: a negotiation, a bail hearing, a job interview, and an auction. In each case, causal relationships are both proposed and tested by the system, finding evidence for some and not others. We provide evidence that the insights from these simulations of social interactions are not available to the LLM purely through direct elicitation. When given its proposed structural causal model for each scenario, the LLM is good at predicting the signs of estimated effects, but it cannot reliably predict the magnitudes of those estimates. In the auction experiment, the in silico simulation results closely match the predictions of auction theory, but elicited predictions of the clearing prices from the LLM are inaccurate. However, the LLM's predictions are dramatically improved if the model can condition on the fitted structural causal model. In short, the LLM knows more than it can (immediately) tell.
연구 동기 및 목표
- SCMs를 청사진으로 사용하여 에이전트를 생성하고 실험을 설계하며 LLMs로 데이터를 분석하는 워크플로를 형식화한다.
- 사회과학 질문에 대한 가설 생성과 인실리코 가설 검정을 자동화한다.
- 여러 시나리오에 걸쳐 접근법을 시연하고 LLM 예측을 이론 및 시뮬레이션 결과와 비교한다.
제안 방법
- 단순 선형 SCMs로 인과관계를 표현하여 가설 생성 및 실험 설계를 유도한다.
- 에이전트를 외생적 SCM 차원에서 변동하는 LLM-전력 엔티티로 구현한다.
- 대화 시뮬레이션과 데이터 수집을 위해 에이전트 간 대화 순서를 교대하는 프로토콜을 사용한다.
- 외생적 차원을 가로지르는 병렬 시뮬레이션을 실행하고 선형 SCM을 추정하여 경로 계수를 얻는다.
- 데이터 분석과 해석을 안내하기 위해 SCM에 사전 분석 계획을 내재한다.
- LLM이 예측한 경로 부호와 크기를 시뮬레이션 추정치 및 이론과 비교한다.

실험 결과
연구 질문
- RQ1SCM-가이드 자율 시스템이 LLM-파워드 에이전트를 사용하여 사회과학 가설을 생성하고 테스트할 수 있는가?
- RQ2인실리코 시뮬레이션이 협상, 보석금 결정, 면접, 경매에서 알려진 이론적 및 경험적 패턴을 재현하는가?
- RQ3LLMs가 효과의 방향과 크기를 예측할 수 있는가, 그리고 적합된 SCM에 조건화될 때 예측이 어떻게 개선되는가?
주요 결과
- 본 시스템은 네 가지 시나리오에 걸쳐 검증 가능한 가설을 생성하고 테스트했으며, 여러 인과 경로에서 유의한 효과를 발견했다.
- 협상 시나리오에서 구매자 예산, 판매자 최저가, 판매자 선호가 거래 확률에 유의하게 영향을 미쳤으며, 크기도 정량화됐다.
- 보석금 결정 시나리오에서 피고의 과거 경력이 보석금을 크게 올렸고, 후회감은 작거나 조건부 효과를 보였다.
- 직무 면접 시나리오에서 합격(bar 합격)이 채용에 큰 양의 영향을 주었고, 키와 면접관의 친근감은 강건한 예측 변수가 아니었다.
- 경매에서 입찰자의 예산이 최종 가격에 양의 영향을 주었으며, 크기는 open-ascension 경매 이론과 일치했다.
- LLM-전용 프롬프트는 경로 계수나 결과를 예측하는 데 시뮬레이션 결과보다 정확도가 떨어졌으며, 다만 적합된 SCM에 조건화할 때 예측이 개선되었다.

더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.