Skip to main content
QUICK REVIEW

[論文レビュー] Beyond accuracy: quantifying trial-by-trial behaviour of CNNs and humans by measuring error consistency

Robert Geirhos, Kristof Meding|arXiv (Cornell University)|Jun 30, 2020
Neural dynamics and brain function被引用数 36
ひとこと要約

試行ごとのエラーチェックの一貫性を用いて人間とCNNの意思決定戦略を比較する方法を提案;CNN同士の一貫性は高いが、人間とCNNの一貫性は偶然レベル、CORnet-Sは人間データよりもフィードフォワードのResNet-50に近い振る舞いを示す。

ABSTRACT

A central problem in cognitive science and behavioural neuroscience as well as in machine learning and artificial intelligence research is to ascertain whether two or more decision makers (be they brains or algorithms) use the same strategy. Accuracy alone cannot distinguish between strategies: two systems may achieve similar accuracy with very different strategies. The need to differentiate beyond accuracy is particularly pressing if two systems are near ceiling performance, like Convolutional Neural Networks (CNNs) and humans on visual object recognition. Here we introduce trial-by-trial error consistency, a quantitative analysis for measuring whether two decision making systems systematically make errors on the same inputs. Making consistent errors on a trial-by-trial basis is a necessary condition for similar processing strategies between decision makers. Our analysis is applicable to compare algorithms with algorithms, humans with humans, and algorithms with humans. When applying error consistency to object recognition we obtain three main findings: (1.) Irrespective of architecture, CNNs are remarkably consistent with one another. (2.) The consistency between CNNs and human observers, however, is little above what can be expected by chance alone -- indicating that humans and CNNs are likely implementing very different strategies. (3.) CORnet-S, a recurrent model termed the "current best model of the primate ventral visual stream", fails to capture essential characteristics of human behavioural data and behaves essentially like a standard purely feedforward ResNet-50 in our analysis. Taken together, error consistency analysis suggests that the strategies used by human and machine vision are still very different -- but we envision our general-purpose error consistency analysis to serve as a fruitful tool for quantifying future progress.

研究の動機と目的

  • 人間とCNNの間で共通の戦略を推測するには、精度だけでは不十分である理由を動機づける。
  • エラ一貫性を、試行ごとの共有エラーの指標として定義し、運用化する。
  • 観測されるエラー重複とコーエンのκの統計的境界と信頼区間をこの文脈で開発する。
  • この手法を用いて、視覚的物体認識課題において複数のCNNと人間、およびCNN同士を比較する。
  • ImageNet精度の向上がより人間のようなエラーパターンにつながるかどうかを評価する。

提案手法

  • 観測されたエラーの重複,c_obsを、同一の正誤応答を示す試行の分率として定義する。
  • 独立した二項決定者を用いて期待重複c_expを、p_i と p_j を用いて計算する(式1)。
  • 独立した観測者のシミュレーションおよび解析的境界(式2–3;付録のセクションS.3)を用いて信頼区間を説明する。
  • コーエンのκでエラー一貫性を定量化する:κ = (c_obs − c_exp) / (1 − c_exp)(式4)。
  • κの境界をc_expの関数として提供する(式5–6)。
  • cue-conflict、edge、silhouette、ImageNet実験からの刺激を用いて比較する;CNN(ImageNet訓練済み)と人間の観察者を分析する。
  • CORnet-Sをリカレントモデルとして含め、ResNet-50のベースラインおよびBrain-Score指標と比較する。

実験結果

リサーチクエスチョン

  • RQ1CNNと人間は同じ刺激上で偶然を超えた一貫した誤りを生み出すか?
  • RQ2より高いImageNet精度はCNNにおいてより人間らしいエラーパターンを意味するか?
  • RQ3アーキテクチャ間および人間と異なるCNNファミリー間でエラーの一貫性はどう変化するか?
  • RQ4CORnet-Sのようなリカレントモデルは、純粋なフィードフォワードモデルと比べて人間とのエラー一貫性を高めるか?
  • RQ5基準となる偶然期待値を前提としたエラー一貫性の限界と境界は何か?

主な発見

  • CNNはアーキテクチャを横断して極めて一貫している。
  • 人間-CNNのエラー一貫性は偶然を少し上回る程度で、人間とCNNの戦略が異なることを示唆する。
  • CORnet-Sは人間とのエラー一貫性でResNet-50にほとんど改善を示さず、この分析ではフィードフォワードモデルのように振る舞う。
  • CNN-CNNのエラー一貫性は通常高く、場合によっては人間対人間の一貫性を上回ることもある。
  • CORnet-SやResNet-50でもCNN-CNNの一貫性は高く、再帰的モデルはこの指標で顕著な挙動パターンを生まない可能性を示す。
  • 分析は、神経予測性や総合的精度を超えた行動指標の評価の重要性を強調する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。