Skip to main content
QUICK REVIEW

[論文レビュー] Learning with Noisy Labels Revisited: A Study Using Real-World Human Annotations

Jiaheng Wei, Zhaowei Zhu|arXiv (Cornell University)|Oct 22, 2021
Machine Learning and Data Classification参考文献 45被引用数 62
ひとこと要約

実世界の人間が注釈したノイズ付きラベルを持つ CIFAR-10N および CIFAR-100N ベンチマークを導入し、インスタンス依存の人間ノイズは合成のクラス依存ノイズとは異なり、ロバストネス評価に影響することを示す。

ABSTRACT

Existing research on learning with noisy labels mainly focuses on synthetic label noise. Synthetic noise, though has clean structures which greatly enabled statistical analyses, often fails to model real-world noise patterns. The recent literature has observed several efforts to offer real-world noisy datasets, yet the existing efforts suffer from two caveats: (1) The lack of ground-truth verification makes it hard to theoretically study the property and treatment of real-world label noise; (2) These efforts are often of large scales, which may result in unfair comparisons of robust methods within reasonable and accessible computation power. To better understand real-world label noise, it is crucial to build controllable and moderate-sized real-world noisy datasets with both ground-truth and noisy labels. This work presents two new benchmark datasets CIFAR-10N, CIFAR-100N, equipping the training datasets of CIFAR-10, CIFAR-100 with human-annotated real-world noisy labels we collected from Amazon Mechanical Turk. We quantitatively and qualitatively show that real-world noisy labels follow an instance-dependent pattern rather than the classically assumed and adopted ones (e.g., class-dependent label noise). We then initiate an effort to benchmarking a subset of the existing solutions using CIFAR-10N and CIFAR-100N. We further proceed to study the memorization of correct and wrong predictions, which further illustrates the difference between human noise and class-dependent synthetic noise. We show indeed the real-world noise patterns impose new and outstanding challenges as compared to synthetic label noise. These observations require us to rethink the treatment of noisy labels, and we hope the availability of these two datasets would facilitate the development and evaluation of future learning with noisy label solutions. Datasets and leaderboards are available at http://noisylabels.com.

研究の動機と目的

  • 合成モデルを超えた実世界のラベルノイズの研究を動機づけ、真値ラベルとノイズラベルを含むアクセス可能なベンチマークを提供する。
  • CIFAR-10およびCIFAR-100における人間が注釈したノイズの分布とパターンを特徴づける。
  • 実際の人間ノイズに対して広範な堅牢学習法をベンチマークし、合成ノイズと比較する。
  • 実世界のノイズ付き監督下での memorization の振る舞いを調査し、学習ダイナミクスを理解する。

提案手法

  • CIFAR-10N と CIFAR-100N を作成するため、CIFAR-10 では画像ごとに3つの人間注釈を Amazon Mechanical Turk で収集し、CIFAR-100 では画像ごとに1つを収集する。評価のため CIFAR の真値ラベルを保持する。
  • ノイズパターンを分析し、データセット間でインスタンス依存・不均衡・特徴量相関のあるラベル反転を示す。
  • ノイズ遷移行列を用いた仮説検定と特徴依存性の評価で、人間注釈ノイズと合成のクラス依存ノイズを比較する。
  • CIFAR-10N および CIFAR-100N で広範な堅牢学習法(損失修正、再加重、正則化、半教師あり風のサンプリング)の評価を行う。
  • 実世界のノイズ下での memorization 振る舞いを調べ、合成ノイズのシナリオとの差異を明らかにする。

実験結果

リサーチクエスチョン

  • RQ1実世界の人間注釈は、クラス依存の合成ノイズとは異なるインスタンス依存のノイズパターンを示すのか。
  • RQ2CIFAR-10N および CIFAR-100N のノイズ分布は、類似クラス間の不均衡およびラベル反転という点でどう異なるか。
  • RQ3一般的な堅牢学習法は、実世界のノイズ付きラベルと合成ノイズのどちらに対して性能を発揮するか。
  • RQ4人間が生成したラベルノイズで学習する場合と合成ノイズの場合で、どのような memorization ダイナミクスが生じるのか。

主な発見

  • 実世界の人間ノイズはインスタンス依存であり、不均衡・特徴量相関のあるラベル反転を示し、合成のクラス依存ノイズだけでは十分に捉えられない。
  • 人間は視覚的に類似するクラス間で誤ラベルをつけやすく、CIFAR-100画像には複数のクリーンラベルが共存することがあり、新しいノイズパターンを生み出す。
  • 方法を問わず、実世界のノイズは合成ノイズより学習上の難しさをもたらし、顕著な性能ギャップと memorization の違いを生む。
  • いくつかの手法(例:ELR+、Divide-Mix)は合成ノイズに対して堅牢な性能を示すが、人間ノイズ下では異なる挙動を示し、ある設定では人間ノイズが特定の手法でわずかに良い性能を促進することさえある。
  • memorization 研究は、ニューラルネットワークが人間ノイズの下で誤ったラベルをより容易に記憶化する傾向があり、合成ノイズとは異なる学習ダイナミクスを示すことを明らかにする。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。