[論文レビュー] Picking on the Same Person: Does Algorithmic Monoculture lead to Outcome Homogenization?
本論文は結果の均質化をアルゴリズム的単一文化のリスクとして定式化し、訓練データの共有とファウンデーションモデルの共有が、公平性ベンチマーク・視覚・言語タスク全体で個人とグループに対する同質的なネガティブな結果へ与える影響を実証的に検証する。
As the scope of machine learning broadens, we observe a recurring theme of algorithmic monoculture: the same systems, or systems that share components (e.g. training data), are deployed by multiple decision-makers. While sharing offers clear advantages (e.g. amortizing costs), does it bear risks? We introduce and formalize one such risk, outcome homogenization: the extent to which particular individuals or groups experience negative outcomes from all decision-makers. If the same individuals or groups exclusively experience undesirable outcomes, this may institutionalize systemic exclusion and reinscribe social hierarchy. To relate algorithmic monoculture and outcome homogenization, we propose the component-sharing hypothesis: if decision-makers share components like training data or specific models, then they will produce more homogeneous outcomes. We test this hypothesis on algorithmic fairness benchmarks, demonstrating that sharing training data reliably exacerbates homogenization, with individual-level effects generally exceeding group-level effects. Further, given the dominant paradigm in AI of foundation models, i.e. models that can be adapted for myriad downstream tasks, we test whether model sharing homogenizes outcomes across tasks. We observe mixed results: we find that for both vision and language settings, the specific methods for adapting a foundation model significantly influence the degree of outcome homogenization. We conclude with philosophical analyses of and societal challenges for outcome homogenization, with an eye towards implications for deployed machine learning systems.
研究の動機と目的
- アルゴリズム的単一文化の下での体系的な harm の一形態として、結果の均質化のリスクを動機づけ、定式化する。
- 個人レベルおよびグループレベルで均質化を測定する数学的枠組みを提案し、実装可能にする。
- ベンチマークを横断してデータ共有とファウンデーションモデル共有を分析し、構成要素共有仮説を実証的に検証する。
- 展開された機械学習システムに対する哲学的・社会的含意を強調する。
提案手法
- 各意思決定者モデル h^i に対して失敗 F^i を定義し、全てのモデルが個人に対して失敗する確率を系統的失敗とする。
- 個人レベルの均質化指標 H^{individual} = 系統的失敗 / (∏_i fail(h^i)) を導入し、均質化と全体の精度を分離する。
- 重み付け方式(平均、等分布、最悪)を用いてグループ均質化 H_{G}^{group} に拡張する。
- タスクとモデルファミリーを跨ぐデータ共有を、固定データ共有と分離データ共有で実験し、共有データ下と非共有データ下の均質化を比較する。
- 視覚領域の scratch、linear probing、finetuning、言語領域の linear probing、finetuning、BitFit などのファウンデーションモデルベースの適応手法を実験し、タスク間の均質化への影響を評価する。
実験結果
リサーチクエスチョン
- RQ1意思決定者間での訓練データ共有は、個人およびグループの結果の均質化を高めるか。
- RQ2ファウンデーションモデルの共有と異なる適応手法は、視覚・言語タスクを横断して結果の均質化を増幅させるのか、それとも緩和するのか。
- RQ3個人レベルとグループレベルの均質化はどのように比較されるか、フェアネス分析にどんな含意があるか。
- RQ4均質化と精度、公平性、頑健性といった従来指標との関係は何か。
- RQ5展開されたシステムにおける結果の均質化から生じる哲学的・社会的課題は何か。
主な発見
- データ共有は結果の均質化を高める;固定共有(同じデータ)では、データセット間・モデル間で分離共有よりも均質化が大きくなる。
- ACS PUMS 実験では個人レベルの均質化がグループレベルを上回り、グループレベルの効果が抑制されているように見える場合でも個人は体系的な害を経験し得ることを示唆する。
- ファウンデーションモデルの共有は混在した結果をもたらす;タスク適応の度合いと機構(プロービング対フィネチューニングなど)が均質化に顕著な影響を与える。
- 視覚タスクでは線形プロービングがフィネチューニングより均質化された結果を生みやすく、言語ではプロービングがフィネチューニング/BitFit よりしばしば均質的。Scratchモデルは一部の視覚設定で最も均質になり得る。
- 特定個人を影響する体系的害を見逃さないよう、個人中心の分析を強く重視すべきである。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。