Skip to main content
QUICK REVIEW

[논문 리뷰] Bias and Fairness in Large Language Models: A Survey

Isabel O. Gallegos, Ryan A. Rossi|arXiv (Cornell University)|2023. 09. 02.
Text Readability and Simplification인용 수 58
한 줄 요약

이 설문조사는 LLM에서의 사회적 편향과 공정성의 정의를 통합하고, 편향 평가 메트릭스와 데이터셋에 대한 분류체계를 제시하며, 사전 처리, 학습 중, 내부 처리, 사후 처리에 걸친 편향 완화 기술을 분류한다.

ABSTRACT

Rapid advancements of large language models (LLMs) have enabled the processing, understanding, and generation of human-like text, with increasing integration into systems that touch our social sphere. Despite this success, these models can learn, perpetuate, and amplify harmful social biases. In this paper, we present a comprehensive survey of bias evaluation and mitigation techniques for LLMs. We first consolidate, formalize, and expand notions of social bias and fairness in natural language processing, defining distinct facets of harm and introducing several desiderata to operationalize fairness for LLMs. We then unify the literature by proposing three intuitive taxonomies, two for bias evaluation, namely metrics and datasets, and one for mitigation. Our first taxonomy of metrics for bias evaluation disambiguates the relationship between metrics and evaluation datasets, and organizes metrics by the different levels at which they operate in a model: embeddings, probabilities, and generated text. Our second taxonomy of datasets for bias evaluation categorizes datasets by their structure as counterfactual inputs or prompts, and identifies the targeted harms and social groups; we also release a consolidation of publicly-available datasets for improved access. Our third taxonomy of techniques for bias mitigation classifies methods by their intervention during pre-processing, in-training, intra-processing, and post-processing, with granular subcategories that elucidate research trends. Finally, we identify open problems and challenges for future work. Synthesizing a wide range of recent research, we aim to provide a clear guide of the existing literature that empowers researchers and practitioners to better understand and prevent the propagation of bias in LLMs.

연구 동기 및 목표

  • NLP 및 LLM에서 사회적 편향 및 공정성 개념을 통합하고 형식화한다.
  • 데이터 구조 및 모델 접근성에 따라 편향 평가 메트릭을 조직하는 분류체계를 개발한다.
  • LLM용 공개 가능한 편향 평가 데이터셋을 수집하고 분류한다.
  • 개입 단계별로 편향 완화 기법을 분류하고 방법에 대한 통일된 표기법을 제공한다.
  • 향후 연구를 안내할 개방된 문제와 도전과제를 식별한다.

제안 방법

  • NLP 및 LLM에 맞춘 LLM 개념과 공정성 요구사항을 형식화한다.
  • (i) 편향 평가 메트릭(임베딩, 확률, 생성 텍스트), (ii) 편향 평가 데이터셋(대체 입력, 프롬프트), (iii) 편향 완화 기법(사전-, 학습 중-, 내부 처리-, 사후 처리)을 제시한다.
  • 메트릭을 비교하고 기법을 형식화하기 위한 통합 수학적 표기법을 제공한다.
  • 편향 평가를 위한 공개 데이터셋을 수집하고 공개한다.
  • LLM의 편향 감소를 위한 개방 문제와 향후 방향을 논의한다.

실험 결과

연구 질문

  • RQ1LLM 및 NLP 작업과 관련된 사회적 편향 및 공정성의 정확한 측면은 무엇인가?
  • RQ2일관된 평가를 가능하게 하도록 데이터 구조와 모델 접근성에 따라 편향 평가 메트릭을 어떻게 조직할 수 있는가?
  • RQ3편향 평가에 사용되는 데이터셋은 무엇이며 표준화하거나 통합할 수 있는 방법은 무엇인가?
  • RQ4개입 단계별로 가장 적합한 편향 완화 기법의 분류체계는 무엇인가?
  • RQ5LLM에서의 공정성 달성을 위한 주요 개방 과제와 향후 방향은 무엇인가?

주요 결과

  • 본 논문은 사회적 편향의 형식적 정의, 그룹 및 개인의 공정성, 그리고 NLP 및 LLM에 적용 가능한 해의 분류(대표적/할당적)를 제공합니다.
  • 임베딩, 확률, 생성 텍스트를 아우르는 편향 평가 메트릭의 통합적 분류체계를 제시하고, 메트릭과 평가 데이터셋 간의 연관성을 명확히 합니다.
  • 구조별로(대체 반사 입력, 프롬프트) 편향 평가 데이터셋을 통합하고, 대상 해와 사회적 집단을 문서화하며, 접근 가능한 공개 저장소를 제공합니다.
  • 개입 단계(사전-, 학습 중-, 내부 처리, 사후 처리)로 구성된 완화 기법의 분류체계를 제시하고, 세부 하위범주와 형식을 제공합니다.
  • 본 설문은 공정성 개념의 강건성, 평가 기준, NLP 생애주기 전반에 걸친 완화 노력의 확장을 포함한 개방 문제를 강조합니다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.