Skip to main content
QUICK REVIEW

[論文レビュー] The Values Encoded in Machine Learning Research

Abeba Birhane, Pratyusha Kalluri|arXiv (Cornell University)|Jun 29, 2021
Big Data and Business Intelligence被引用数 12
ひとこと要約

本論文は、ICMLおよびNeurIPSの100編の高いインパクトを持つ機械学習研究論文に埋め込まれた価値を特定・分析するための新規アノテーションスキームを導入する。その結果、パフォーマンス、一般化、独創性といったトップの価値が、社会的ニーズや負の影響にほとんど注意を払わず、エリート機関やテック企業と強く結びついた形で体系的に定義されており、権力の集中を助長する傾向にあることが明らかになった。

ABSTRACT

Machine learning currently exerts an outsized influence on the world, increasingly affecting institutional practices and impacted communities. It is therefore critical that we question vague conceptions of the field as value-neutral or universally beneficial, and investigate what specific values the field is advancing. In this paper, we first introduce a method and annotation scheme for studying the values encoded in documents such as research papers. Applying the scheme, we analyze 100 highly cited machine learning papers published at premier machine learning conferences, ICML and NeurIPS. We annotate key features of papers which reveal their values: their justification for their choice of project, which attributes of their project they uplift, their consideration of potential negative consequences, and their institutional affiliations and funding sources. We find that few of the papers justify how their project connects to a societal need (15\%) and far fewer discuss negative potential (1\%). Through line-by-line content analysis, we identify 59 values that are uplifted in ML research, and, of these, we find that the papers most frequently justify and assess themselves based on Performance, Generalization, Quantitative evidence, Efficiency, Building on past work, and Novelty. We present extensive textual evidence and identify key themes in the definitions and operationalization of these values. Notably, we find systematic textual evidence that these top values are being defined and applied with assumptions and implications generally supporting the centralization of power.Finally, we find increasingly close ties between these highly cited papers and tech companies and elite universities.

研究の動機と目的

  • 影響力のあるML研究に埋め込まれた具体的な価値を特定することで、機械学習が価値中立的であるという神話に挑戦すること。
  • 研究論文に埋め込まれた価値を検出可能な細分化されたアノテーションスキームの開発、特に根拠、価値の上昇、リスクの検討を含む。
  • ICMLおよびNeurIPSの100編の高いインパクトを持つ論文を分析し、支配的価値とそれらの技術的議論における実装方法を解明すること。
  • これらの価値が機関の所属関係、資金源、権力の非対称性によってどのように形成されているかを調査すること。
  • パフォーマンスや独創性といった価値が、AI研究における既存の権力構造を強化する形で定義・適用されている仕組みを明らかにすること。

提案手法

  • 研究論文に埋め込まれた価値を体系的にラベル付けするためのカスタムアノテーションスキームを開発。プロジェクトの根拠、価値の上昇、リスクの検討、機関的文脈に焦点を当てる。
  • 2018年から2021年までのICMLおよびNeurIPSの100編の高いインパクトを持つ論文にこのスキームを適用し、3,500句を超えるアノテーションを取得。
  • 文書ごとの線形的テクスト分析を実施し、論文に現れる59の異なる価値を特定・分類。
  • 特にパフォーマンス、一般化、効率性、過去の研究の積み重ね、独創性といった価値の繰り返しのコミットメントを訓練と抽出。
  • 機関の所属関係と資金源をマッピングし、エリート大学とテック企業との間の関係を追跡。
  • 定性的および定量的分析を実施し、価値の定義と実装方法、特に権力の集中と関連して分析。

実験結果

リサーチクエスチョン

  • RQ1高いインパクトを持つ機械学習研究論文に埋め込まれた価値は何か。それらは技術的議論の中でどのように実装されているか。
  • RQ2これらの論文が社会的ニーズや潜在的な負の影響に基づいてプロジェクトを正当化する割合はどの程度か。
  • RQ3機関の所属関係や資金源は、ML研究で強調される価値とどの程度相関しているか。
  • RQ4パフォーマンスや独創性といったコア価値が、どのように権力の集中を助長する形で定義・適用されているか。
  • RQ5上位のML会議で提唱される価値は、システム的な権力の非対称性をどのように反映または再現しているか。

主な発見

  • 100編の論文のうち15%しか、社会的ニーズに基づいたプロジェクトの正当化を行っておらず、現実の社会的関連性にほとんど注目されていないことが示された。
  • 論文の1%しか、潜在的な負の影響について言及しておらず、倫理的予見性と影響評価の大きな空白が明らかになった。
  • 最も頻繁に正当化された5つの価値(パフォーマンス、一般化、定量的証拠、効率性、独創性)は、中央集権的な権力構造を好む形で体系的に定義されていた。
  • 独創性やパフォーマンスといった価値が、既存の機関やテック企業に利益をもたらす仮定に基づいて実装されているという明確なテクスト的証拠が得られた。
  • 分析から、上位のML論文とエリート大学、主要なテック企業との間で、ますます密接な機関的・資金的つながりが確認された。
  • ML研究に59の異なる価値が同定されたが、上位5つの価値は主に技術的ではあるが、その適用と含みにおいて社会的・政治的に重要な意味を持つことが明らかになった。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。