[논문 리뷰] Beyond accuracy: quantifying trial-by-trial behaviour of CNNs and humans by measuring error consistency
트라이얼-별 오류 일관성을 사용하여 인간과 CNN 간 의사결정 전략을 비교하는 방법을 도입; CNN은 서로 매우 일관적이지만 인간- CNN은 우연 수준의 일관성만 보이며 CORnet-S는 인간 데이터보다 전방향 ResNet-50처럼 작동한다.
A central problem in cognitive science and behavioural neuroscience as well as in machine learning and artificial intelligence research is to ascertain whether two or more decision makers (be they brains or algorithms) use the same strategy. Accuracy alone cannot distinguish between strategies: two systems may achieve similar accuracy with very different strategies. The need to differentiate beyond accuracy is particularly pressing if two systems are near ceiling performance, like Convolutional Neural Networks (CNNs) and humans on visual object recognition. Here we introduce trial-by-trial error consistency, a quantitative analysis for measuring whether two decision making systems systematically make errors on the same inputs. Making consistent errors on a trial-by-trial basis is a necessary condition for similar processing strategies between decision makers. Our analysis is applicable to compare algorithms with algorithms, humans with humans, and algorithms with humans. When applying error consistency to object recognition we obtain three main findings: (1.) Irrespective of architecture, CNNs are remarkably consistent with one another. (2.) The consistency between CNNs and human observers, however, is little above what can be expected by chance alone -- indicating that humans and CNNs are likely implementing very different strategies. (3.) CORnet-S, a recurrent model termed the "current best model of the primate ventral visual stream", fails to capture essential characteristics of human behavioural data and behaves essentially like a standard purely feedforward ResNet-50 in our analysis. Taken together, error consistency analysis suggests that the strategies used by human and machine vision are still very different -- but we envision our general-purpose error consistency analysis to serve as a fruitful tool for quantifying future progress.
연구 동기 및 목표
- 왜 정확도만으로 인간과 CNN 간 공유 전략을 추론하기에 충분하지 않은지의 동기를 제시한다.
- 오류 일관성을 공유된 오류의 트라이얼-별 측정으로 정의하고 운용화한다.
- 관찰된 오류 중복 및 Cohen’s kappa에 대한 통계적 경계값과 신뢰구간을 이 맥락에서 개발한다.
- 시각적 대상 인식 작업에서 여러 CNN과 인간, 그리고 서로 간의 비교에 방법을 적용한다.
- ImageNet 정확도 향상이 인간에 더 유사한 오류 패턴으로 이어지는지 평가한다.
제안 방법
- 관찰된 오류 중복 c_obs를 동일한 정답/오답 응답의 트라이얼 비율로 정의한다.
- p_i와 p_j를 사용하여 독립 이항 의사결정자 하에서 예상 중복 c_exp를 계산한다( Eq. 1 ).
- 독립 관찰자 시뮬레이션 및 해석적 경계( Eq. 2–3; 부록의 섹션 S.3 )를 통해 신뢰구간을 기술한다.
- Cohen의 kappa를 사용하여 오류 일관성을 정량화한다: kappa = (c_obs − c_exp) / (1 − c_exp) (Eq. 4).
- c_exp의 함수로서의 kappa에 대한 경계값을 제공한다( Eq. 5–6 ).
- cue-conflict, edge, silhouette, 및 ImageNet 실험에서의 자극을 사용하고 CNN( ImageNet- trained )과 인간 관찰자를 비교한다.
- CORnet-S를 순환 모델로 포함하고 기반선 ResNet-50 및 Brain-Score 지표와 비교한다.
실험 결과
연구 질문
- RQ1CNN과 인간이 같은 자극에서 우연을 넘어서 일관된 오류를 생성하는가?
- RQ2더 높은 ImageNet 정확도가 CNN에서 더 인간에 가까운 오류 패턴을 시사하는가?
- RQ3아키텍처 간 및 인간과 서로 다른 CNN 계통 간 오류 일관성은 어떻게 variation을 보이는가?
- RQ4CORnet-S와 같은 순환 모델이 순방향 모델에 비해 인간과의 오류 일관성을 더 높이는가?
- RQ5기저 우연 기대치에 비해 오류 일관성의 한계와 경계는 무엇인가?
주요 결과
- CNN은 아키텍처 전반에 걸쳐 서로 놀랍게도 일관성이 높다.
- 인간- CNN 오류 일관성은 우연 수준을 약간 상회하는 수준으로 인간과 CNN 간 전략이 다름을 시사한다.
- CORnet-S는 인간과의 오류 일관성에서 ResNet-50에 비해 의미 있는 개선을 보이지 않으며 이 분석에서 순방향 모델처럼 작동한다.
- CNN-CNN 오류 일관성은 일반적으로 높으며 때로는 인간-인간 일관성보다도 큰 경우가 있다.
- 심지어 CORnet-S와 ResNet-50도 CNN-CNN 일관성이 높게 나타나 순환 모델이 이 지표에서 뚜렷한 행동 패턴 차이를 제공하지 못할 수 있음을 시사한다.
- 분석은 신경 예측력이나 전체 정확도 외의 행동 지표 평가의 중요성을 강조한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.