Skip to main content
QUICK REVIEW

[논문 리뷰] Bias Dynamics in BabyLMs: Towards a Compute-Efficient Sandbox for Democratising Pre-Training Debiasing

Filip Trhlík, Andrew Caines|arXiv (Cornell University)|2026. 01. 14.
Artificial Intelligence in Healthcare and Education인용 수 0
한 줄 요약

본 논문은 BabyLM이 표준 LMs의 편향 획득 및 편향 제거 역학을 재현하고, 사전 학습 편향 제거 실험을 위한 비용 효율적 샌드박스로 활용될 수 있으며 계산 자원을 500 GPU시간에서 약 30 GPU시간 수준으로 줄일 수 있음을 보여준다.

ABSTRACT

Pre-trained language models (LMs) have, over the last few years, grown substantially in both societal adoption and training costs. This rapid growth in size has constrained progress in understanding and mitigating their biases. Since re-training LMs is prohibitively expensive, most debiasing work has focused on post-hoc or masking-based strategies, which often fail to address the underlying causes of bias. In this work, we seek to democratise pre-model debiasing research by using low-cost proxy models. Specifically, we investigate BabyLMs, compact BERT-like models trained on small and mutable corpora that can approximate bias acquisition and learning dynamics of larger models. We show that BabyLMs display closely aligned patterns of intrinsic bias formation and performance development compared to standard BERT models, despite their drastically reduced size. Furthermore, correlations between BabyLMs and BERT hold across multiple intra-model and post-model debiasing methods. Leveraging these similarities, we conduct pre-model debiasing experiments with BabyLMs, replicating prior findings and presenting new insights regarding the influence of gender imbalance and toxicity on bias formation. Our results demonstrate that BabyLMs can serve as an effective sandbox for large-scale LMs, reducing pre-training costs from over 500 GPU-hours to under 30 GPU-hours. This provides a way to democratise pre-model debiasing research and enables faster, more accessible exploration of methods for building fairer LMs.

연구 동기 및 목표

  • LM 편향 형성 및 편향 제거 연구를 위한 저비용 샌드박스 도입의 필요성 제기.
  • BabyLM이 표준 LMs와 유사하게 편향을 획득하고 편향 제거 방법에 유사한 방식으로 반응함을 보인다.
  • 사전 모델 편향 제거 실험을 상당히 축소된 계산 자원으로 실행할 수 있음을 보여준다.

제안 방법

  • BabyLM LTG-BERT 변형을 사용하고 표준 BERT와 비교하여 편향 및 성능 탐색 지표(BLiMP, EWoK, CrowS-Pairs, StereoSet)를 평가한다.
  • 여러 탐침의 점수를 평균하여 합성 편향 지표와 합성 성능 지표를 설정한다.
  • 합성 성능과 편향 간의 상관관계를 분석하여 BabyLM을 표준 LMs의 대리 지표로서 검증한다(표 1).
  • 사후-모델(post-model) 및 intra-model 방법(Sent-Debias, INLP, CDA, CDS, debiasing losses, dropout)을 사용하여 편향 제거 변화의 상관관계를 평가한다.
  • LTG-Baseline에서 사전 모델 편향 제거 실험(CDA, 독성 제거, 변동 증강)을 수행하여 비용과 효과를 평가한다.

실험 결과

연구 질문

  • RQ1BabyLM이 BERT와 같은 더 큰 LMs에서 관찰된 편향 획득 역학을 재현할 수 있는가?
  • RQ2사후-모델, 내부-모델, 사전-모델 개입 전반에 걸쳐 BabyLM이 표준 LM과 비교할 만한 편향 제거 행태를 보이는가?
  • RQ3큰 규모의 실험에 착수하기 전에 사전 모델 편향 제거 전략을 탐구하기 위한 비용 효율적 플랫폼으로 BabyLM이 역할을 할 수 있는가?

주요 결과

  • BabyLM은 합성 성능과 합성 편향 간의 강한 양의 상관관계를 보이며 표준 LMs와 유사하다(표 1에서 BabyLM r = 0.833, Standard r = 0.753).
  • 사후-모델 편향 제거 효과는 모델 간에 일관되며, 성별 중심 INLP는 편향을 감소시키는 반면 인종 중심 INLP는 정확도를 해칠 수 있다.
  • 내부-모델 편향 제거는 모델 간에 유사한 편향 감소를 보이며, LTG-Baseline은 성능-편향 변화에서 BERT와 가장 가깝게 정준 상관관계(canonical correlations)를 보인다.
  • BabyLM의 사전 모델 실험(CDA, 독성 제거, 변동 증강)은 알려진 편향 제거 효과를 재현하고 매 실행당 약 30 GPU시간 수준에서 새로운 블레이드를 가능하게 한다.
  • 독성은 다운스트림 편향과 연관되며, 독성 문장을 제거하는 것이 무작위 말뭉치 감소보다 편향을 더 효과적으로 줄인다.
  • BabyLM은 상당히 낮은 계산으로도 기존 결과를 재현하고 새로운 통찰을 제공하므로, 편향 제거 샌드박스로서 실현 가능성이 있음을 시사한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.