[논문 리뷰] Picking on the Same Person: Does Algorithmic Monoculture lead to Outcome Homogenization?
본 논문은 outcome homogenization를 알고리즘적 monoculture의 위험으로 형식화하고, 훈련 데이터 공유와 foundation models의 공유가 개인 및 그룹 전반에 걸친 공정성 벤치마크, 비전 및 언어 과제에서의 동질적인 부정적 결과에 어떤 영향을 미치는지 데이터로 검증한다.
As the scope of machine learning broadens, we observe a recurring theme of algorithmic monoculture: the same systems, or systems that share components (e.g. training data), are deployed by multiple decision-makers. While sharing offers clear advantages (e.g. amortizing costs), does it bear risks? We introduce and formalize one such risk, outcome homogenization: the extent to which particular individuals or groups experience negative outcomes from all decision-makers. If the same individuals or groups exclusively experience undesirable outcomes, this may institutionalize systemic exclusion and reinscribe social hierarchy. To relate algorithmic monoculture and outcome homogenization, we propose the component-sharing hypothesis: if decision-makers share components like training data or specific models, then they will produce more homogeneous outcomes. We test this hypothesis on algorithmic fairness benchmarks, demonstrating that sharing training data reliably exacerbates homogenization, with individual-level effects generally exceeding group-level effects. Further, given the dominant paradigm in AI of foundation models, i.e. models that can be adapted for myriad downstream tasks, we test whether model sharing homogenizes outcomes across tasks. We observe mixed results: we find that for both vision and language settings, the specific methods for adapting a foundation model significantly influence the degree of outcome homogenization. We conclude with philosophical analyses of and societal challenges for outcome homogenization, with an eye towards implications for deployed machine learning systems.
연구 동기 및 목표
- algorithmic monoculture 하에서 outcome homogenization의 위험을 체계적 피해의 한 형태로 동기 부여하고 형식화한다.
- 개인 및 그룹 차원에서 동질화를 측정하기 위한 수학적 프레임워크를 제안하고 이를 운용화한다.
- 벤치마크 전반에 걸친 데이터 공유와 foundation-model 공유를 분석하여 component-sharing 가설을 실증적으로 검증한다.
- 배포된 머신러닝 시스템에 대한 철학적 및 사회적 함의를 강조한다.
제안 방법
- 각 의사결정자 모델 h^i에 대해 실패 F^i를 정의하고, 모든 모델이 한 개인에서 실패하는 확률로 시스템적 실패를 정의한다.
- 개인 동질화 지표 H^{individual} = systemic failure / (prod fail(h^i))를 도입하여 동질화를 전반적 정확도와 구분한다.
- 가중치 체계(평균, 균일, 최악)를 사용하여 그룹 동질화 H_{G}^{group}으로 확장한다.
- 고정된 데이터 공유와 서로 다른 학습 데이터(disjoint training data)를 사용해 다양한 과제 및 모델 계열에서 데이터를 공유한 경우와 공유하지 않은 경우의 동질화를 비교하기 위해 실험한다.
- 비전: scratch, linear probing, finetuning; 언어: linear probing, finetuning, BitFit와 같은 foundation-model 기반 적응 방법을 실험하여 과제 전반에서의 동질화에 미치는 영향을 평가한다.
실험 결과
연구 질문
- RQ1의사결정자 간의 학습 데이터 공유가 개인 및 그룹의 outcome homogenization을 증가시키는가?
- RQ2foundation-model 공유와 다양한 적응 방법이 비전 및 언어 과제에서 outcome homogenization을 증폭시키거나 완화시키는가?
- RQ3개인 차원의 동질화와 그룹 차원의 동질화는 어떻게 비교되며, 공정성 분석에 어떤 함의가 있는가?
- RQ4동질화와 정확도, 공정성, 강건성 같은 전통적 지표 간의 관계는 무엇인가?
- RQ5배포된 시스템에서의 outcome homogenization으로부터 어떤 철학적 및 사회적 문제가 제기되는가?
주요 결과
- 데이터 공유는 outcome homogenization을 증가시킨다; 고정 공유(같은 데이터)가 서로 다른 데이터 세트와 모델에서의 비공유 공유보다 더 많은 동질화를 초래한다.
- ACS PUMS 실험에서 개인 차원의 동질화가 그룹 차원의 동질화를 능가하여, 그룹 수준의 효과가 완만해 보일 때도 개인은 체계적 피해를 경험할 수 있음을 시사한다.
- Foundation-model 공유는 혼재된 결과를 낳으며, 작업 적응의 정도와 메커니즘(예: probing 대 finetuning)이 동질화에 유의하게 영향을 준다.
- 비전 과제에서 선형 프로빙(linear probing)은 일반적으로 finetuning보다 더 동질화된 결과를 나타내고, 언어 과제의 경우 프로빙이 종종 finetuning/BitFit보다 더 동질적이며, scratch 모델이 일부 비전 설정에서 가장 동질화될 수 있다.
- 특정 개인에게 영향을 미치는 체계적 피해를 간과할 수 있으므로 개인 중심 분석에 강한 중점을 두는 것이 필요하다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.