Skip to main content
QUICK REVIEW

[Paper Review] Bias Amplification: Large Language Models as Increasingly Biased Media

Ze Wang, Zekun Wu|arXiv (Cornell University)|Oct 19, 2024
Natural Language Processing Techniques4 citations
TL;DR

This paper introduces a theoretical framework for bias amplification in large language models (LLMs), demonstrating that models can increasingly amplify pre-existing political biases—such as right-leaning tendencies in GPT-2—through iterative fine-tuning on synthetic data, even without model collapse. The study identifies distinct neuronal mechanisms for bias amplification and model collapse, and finds that preservation and accumulation strategies effectively mitigate bias amplification.

ABSTRACT

Model collapse, a phenomenon characterized by performance degradation due to iterative training on synthetic data, has been widely studied. However, its implications for bias amplification, the progressive intensification of pre-existing societal biases in Large Language Models (LLMs), remain significantly underexplored, despite the growing influence of LLMs in shaping online discourse. In this paper, we introduce a open, generational, and long-context benchmark specifically designed to measure political bias amplification in LLMs, leveraging sentence continuation tasks derived from a comprehensive dataset of U.S. political news. Our empirical study using GPT-2 reveals consistent and substantial political bias intensification (e.g., right-leaning amplification) over iterative synthetic training cycles. We evaluate three mitigation strategies, Overfitting, Preservation, and Accumulation, and demonstrate that bias amplification persists independently of model collapse, even when the latter is effectively controlled. Furthermore, we propose a mechanistic analysis approach that identifies neurons correlated with specific phenomena during inference through regression and statistical tests. This analysis uncovers largely distinct neuron populations driving bias amplification and model collapse, underscoring fundamentally different underlying mechanisms. Finally, we supplement our empirical findings with theoretical intuition that explains the separate origins of these phenomena, guiding targeted strategies for bias mitigation.

Motivation & Objective

  • To address the lack of theoretical and empirical understanding of bias amplification in LLMs, distinct from model collapse.
  • To investigate whether LLMs amplify political bias during self-consuming training loops using synthetic data.
  • To evaluate mitigation strategies such as overfitting, preservation, and accumulation for reducing bias amplification.
  • To identify and distinguish the neuronal mechanisms driving bias amplification versus model collapse in LLMs.

Proposed method

  • Proposes a theoretical framework based on weighted maximum likelihood estimation to define necessary and sufficient conditions for bias amplification, independent of model collapse.
  • Employs statistical simulations using weighted maximum likelihood estimation to demonstrate bias amplification without sampling or functional form issues.
  • Develops a high-accuracy political bias classifier to benchmark political leaning in long-text generations, enabling evaluation of bias in open-ended tasks.
  • Conducts iterative fine-tuning of GPT-2 on synthetic data generated by prior iterations to empirically observe increasing right-leaning bias.
  • Applies a novel mechanistic interpretation pipeline to identify neuron-level contributions to bias amplification and model collapse using regression analysis on weight changes.
  • Uses Newey-West standard errors and Bonferroni correction in regression models to test the significance of neuron-weight changes on bias shifts.

Experimental results

Research questions

  • RQ1What are the necessary and sufficient conditions for bias amplification in LLMs, and how can it be theoretically distinguished from model collapse?
  • RQ2To what extent does iterative fine-tuning on synthetic data amplify political bias in GPT-2, and does this bias increase over generations?
  • RQ3How effective are overfitting, preservation, and accumulation strategies in mitigating bias amplification and model collapse?
  • RQ4Are the neuronal mechanisms underlying bias amplification and model collapse distinct, and can they be identified through mechanistic interpretation?

Key findings

  • Bias amplification occurs independently of model collapse, as demonstrated by theoretical analysis and statistical simulations using weighted maximum likelihood estimation.
  • GPT-2 exhibits a progressive increase in right-leaning political bias in sentence continuation tasks after iterative fine-tuning on synthetic data generated by previous iterations.
  • Preservation and accumulation strategies effectively mitigate both bias amplification and model collapse, while overfitting shows limited effectiveness.
  • Mechanistic interpretation reveals minimal overlap between neuron sets responsible for bias amplification and model collapse, supporting the theoretical distinction between the two phenomena.
  • The regression-based neuron analysis identifies specific neurons whose weight changes significantly correlate with shifts in political bias, with statistical significance confirmed via Newey-West standard errors and Bonferroni correction.
  • The study confirms that bias amplification can occur even in the absence of biased training data, indicating that model dynamics themselves can drive bias escalation.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.