[논문 리뷰] Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs
논문은 socio-demographic 페르소나를 LLM에 할당하면 데이터셋과 모델 전반에서 상당한 추론 편향이 유도되며, 명시적 기피와 암묵적 오류 패턴이 모두 나타나고, 간단한 편향 제거 프롬프트는 거의 효과가 없다는 것을 보여준다.
Recent works have showcased the ability of LLMs to embody diverse personas in their responses, exemplified by prompts like 'You are Yoda. Explain the Theory of Relativity.' While this ability allows personalization of LLMs and enables human behavior simulation, its effect on LLMs' capabilities remains unclear. To fill this gap, we present the first extensive study of the unintended side-effects of persona assignment on the ability of LLMs to perform basic reasoning tasks. Our study covers 24 reasoning datasets, 4 LLMs, and 19 diverse personas (e.g. an Asian person) spanning 5 socio-demographic groups. Our experiments unveil that LLMs harbor deep rooted bias against various socio-demographics underneath a veneer of fairness. While they overtly reject stereotypes when explicitly asked ('Are Black people less skilled at mathematics?'), they manifest stereotypical and erroneous presumptions when asked to answer questions while adopting a persona. These can be observed as abstentions in responses, e.g., 'As a Black person, I can't answer this question as it requires math knowledge', and generally result in a substantial performance drop. Our experiments with ChatGPT-3.5 show that this bias is ubiquitous - 80% of our personas demonstrate bias; it is significant - some datasets show performance drops of 70%+; and can be especially harmful for certain groups - some personas suffer statistically significant drops on 80%+ of the datasets. Overall, all 4 LLMs exhibit this bias to varying extents, with GPT-4-Turbo showing the least but still a problematic amount of bias (evident in 42% of the personas). Further analysis shows that these persona-induced errors can be hard-to-discern and hard-to-avoid. Our findings serve as a cautionary tale that the practice of assigning personas to LLMs - a trend on the rise - can surface their deep-rooted biases and have unforeseeable and detrimental side-effects.
연구 동기 및 목표
- 다양한 과제에서 페르소나 할당이 LLM의 추론 능력에 영향을 미치는지 조사한다.
- 24개의 추론 데이터셋에서 19개의 socio-demographic 페르소나와 관련된 편향을 정량화한다.
- 편향이 어떻게 나타나는지(명시적 기피 대 암묵적 오류) 및 모델과 데이터셋 간의 변동성을 서술한다.
- 페르소나 유도 편향을 완화하기 위한 프롬프트 기반 편향 제거 전략과 그 효과를 평가한다.
제안 방법
- 시스템 프롬프트를 통해 네 가지 LLM(ChatGPT-3.5 변형, GPT-4-Turbo, Llama-2-70b-chat)에 페르소나를 할당한다.
- 수학, 법률, 의학, 윤리 등 다양한 영역을 포괄하는 24개의 추론 데이터셋에서 평가한다.
- 5개 사회인구학 그룹에 걸친 19개의 페르소나를 사용하고 세 가지 페르소나 지시 variante로 제로샷 프롬프트를 수행한다.
- Wilson 신뢰구간을 사용하여 인간 기준선 및 평균 인간 기준선에 대한 통계적 유의 차이를 측정한다.
- 공유된 비기피 문제에 대한 성능 및 데이터셋 범주 간 비교를 통해 명시적 기피와 암묵적 편향을 분석한다.
- 해독 변동성을 고려하여 페르소나/데이터셋 쌍당 3회 실행의 결과를 평균화하여 보고한다.
실험 결과
연구 질문
- RQ1다양한 데이터셋에서 페르소나 할당이 LLM 추론의 성능 차이를 야기하는가?
- RQ2사회인구학적 차원에서 페르소나 유도 편향은 얼마나 널리 퍼져 있으며 모델 및 데이터셋에 따라 어떻게 달라지는가?
- RQ3이러한 편향은 어떤 형태로 나타나며(명시적 기피 대 암묵적 오류) 얼마나 탐지 가능한가?
- RQ4단순 프롬프트 기반 편향 제거가 페르소나 유도 편향을 완화할 수 있는가, 그리고 한계는 무엇인가?
- RQ5페르소나 쌍 간에 도메인 또는 과제 특이적 편향 패턴이 존재하는가?
주요 결과
- ChatGPT-3.5 페르소나의 80%에서 데이터셋 전반에 걸친 편향이 나타났고, 일부 데이터셋에서는 정확도에서 상대적 하락이 최대 70%에 이르렀다.
- GPT-4-Turbo는 편향이 가장 적게 나타났으나 여전히 42%의 페르소나에 영향을 보였다.
- Phys. Disabled 및 Religious 페르소나는 평균 정확도 하락이 35% 이상이고 특정 데이터셋에서 최대 69%에 달하는 경우가 흔했다.
- 편향은 모델, 페르소나, 도메인 전반에 만연하며 같은 그룹 내 및 그룹 간 차이가 명확하다(예: Religion 또는 Disability 그룹 내 차이).
- 기피는 많은 오류의 원인으로 작용한다(예: Phys. Disabled 오류의 58%; Atheist 대 Religious의 35% 등), 그러나 암묵적 편향도 비기피 오류를 초래한다.
- Debiasing prompts like don’t refuse or treat human are largely ineffective; task-specific expertise can reduce bias but has limited generalizability.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.