[논문 리뷰] Diminished Diversity-of-Thought in a Standard Large Language Model
이 논문은 사회과학 재현에서 인간 참여자를 대리하는 지표로 GPT-3.5를 테스트하고, 특정 프롬프트가 응답의 차이가 거의 없는 현상을 문서화하여 LLM을 인간 주제의 일반적 대체로 삼는 타당성에 도전한다.
We test whether Large Language Models (LLMs) can be used to simulate human participants in social-science studies. To do this, we run replications of 14 studies from the Many Labs 2 replication project with OpenAI's text-davinci-003 model, colloquially known as GPT3.5. Based on our pre-registered analyses, we find that among the eight studies we could analyse, our GPT sample replicated 37.5% of the original results and 37.5% of the Many Labs 2 results. However, we were unable to analyse the remaining six studies due to an unexpected phenomenon we call the "correct answer" effect. Different runs of GPT3.5 answered nuanced questions probing political orientation, economic preference, judgement, and moral philosophy with zero or near-zero variation in responses: with the supposedly "correct answer." In one exploratory follow-up study, we found that a "correct answer" was robust to changing the demographic details that precede the prompt. In another, we found that most but not all "correct answers" were robust to changing the order of answer choices. One of our most striking findings occurred in our replication of the Moral Foundations Theory survey results, where we found GPT3.5 identifying as a political conservative in 99.6% of the cases, and as a liberal in 99.3% of the cases in the reverse-order condition. However, both self-reported 'GPT conservatives' and 'GPT liberals' showed right-leaning moral foundations. Our results cast doubts on the validity of using LLMs as a general replacement for human participants in the social sciences. Our results also raise concerns that a hypothetical AI-led future may be subject to a diminished diversity-of-thought.
연구 동기 및 목표
- 표준 LLM(GPT-3.5)가 사회과학 재현 연구에서 인간 참여자를 시뮬레이션할 수 있는지 평가합니다.
- 다양한 과제에 걸쳐 Many Labs 2 프로젝트 대비 재현 성공 여부를 측정합니다.
- LLM 응답의 사고 다양성에 해를 끼치는 현상을 식별하고 특성화합니다.
제안 방법
- OpenAI GPT-3.5(text-davinci-003)를 사용하여 14개의 Many Labs 2 연구를 재현합니다.
- 프리레지스트리 분석을 통해 GPT 출력과 원본 및 Many Labs 2 결과를 비교합니다.
- 여덟 개의 분석 가능한 연구를 분석하고 재현 비율을 보고하며 응답의 예기치 않은 거의 0의 차이를 문서화합니다.
- 프롬프트 순서를 포함한 인구통계적 세부 정보의 변화에 따른 탐색적 후속 연구를 수행하여 응답의 견고성을 테스트합니다.
- 다양한 프롬프트 조건에서 도덕적 기초 이론 설문 결과를 검토합니다.
실험 결과
연구 질문
- RQ1GPT-3.5가 원래 Many Labs 2 결과의 상당 부분을 재현할 수 있는가?
- RQ2GPT-3.5 응답이 인간 참여자를 위한 유효한 대리자라고 간주될 만큼 충분한 다양성을 보여주는가?
- RQ3LLMs를 사회과학 재현에 사용하는 데 한계를 만드는 현상(예: '정답' 효과)은 무엇인가?
- RQ4인구통계 및 응답 순서 프롬프트의 변화에 대해 GPT-3.5에서 도출된 결론은 얼마나 견고한가?
- RQ5다양한 사고의 다양성에 관한 AI 주도 미래 시나리오에 대한 시사점은 무엇인가?
주요 결과
- GPT-3.5가 여덟 개의 분석 가능한 연구에서 원래 결과의 37.5%를 재현했습니다.
- GPT-3.5가 Many Labs 2 결과의 37.5%를 재현했습니다.
- 정확한 답을 설명하는 강력한 '정답' 효과가 등장하여 미세한 질문에 대해 응답의 차이가 거의 0이거나 0에 가까웠습니다.
- 탐색적 후속 연구에서 '정답'은 프롬프트 이전의 인구통계 변형에 대해 견고했습니다.
- 도덕적 기초 이론 재현에서 GPT-3.5는 역순의 경우 99.6%에서 보수적이고 99.3%에서 진보적이라고 식별되었으나 두 그룹 모두 오른쪽으로 기울어진 도덕적 기초를 보였습니다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.