[논문 리뷰] Computing a human-like reaction time metric from stable recurrent vision models
논문은 evidential deep learning으로 학습된 안정적 순환 비전 모델에서 파생된 자극-계산 가능한 반응 시간 대리 수 xi_cRNN을 도입하고, 네 가지 시각 과제에서 인간의 RT 패턴과의 정합성을 보인다.
The meteoric rise in the adoption of deep neural networks as computational models of vision has inspired efforts to "align" these models with humans. One dimension of interest for alignment includes behavioral choices, but moving beyond characterizing choice patterns to capturing temporal aspects of visual decision-making has been challenging. Here, we sketch a general-purpose methodology to construct computational accounts of reaction times from a stimulus-computable, task-optimized model. Specifically, we introduce a novel metric leveraging insights from subjective logic theory summarizing evidence accumulation in recurrent vision models. We demonstrate that our metric aligns with patterns of human reaction times for stimulus manipulations across four disparate visual decision-making tasks spanning perceptual grouping, mental simulation, and scene categorization. This work paves the way for exploring the temporal alignment of model and human visual strategies in the context of various other cognitive tasks toward generating testable hypotheses for neuroscience. Links to the code and data can be found on the project page: https://serre-lab.github.io/rnn_rts_site.
연구 동기 및 목표
- 신경망 다이나믹스와 인간의 시각 의사결정 사이의 시간적 정렬을 동기화한다.
- cRNN으로부터 모델 주도적이고 자극-계산 가능한 반응 시간 지표를 개발한다.
- 다양한 과제에 걸쳐 반응 시간 지표가 인간 RT 패턴과 질적으로 정렬됨을 시연한다.
- 신경과학 가설을 생성하고 시간적 다이나믹스를 연구하기 위한 프레임워크를 제공한다.
제안 방법
- C-RBP와 Evidential Deep Learning (EDL)으로 안정적 순환 비전 모델 (cRNN)을 학습하여 클래스에 대한 디리슐 분포 신념을 얻는다.
- RT 지표 xi_cRNN를 시간에 따른 모델 불확실성 곡선의 면적으로 정의하고, xi_cRNN = integral_0^T Epsilon(t) dt라고 한다.
- 끌어당김(애터트) 다이나믹스와 EDL을 사용하여 추가 감독 없이 시간에 따라 변하는 불확실성을 얻는다.
- 인간의 RT와의 정합성을 평가하기 위해 네 가지 과제에 프레임워크를 적용한다: 증가적 그룹화, 시각적 시뮬레이션(Planko), 미로 경로 추론, 장면 분류.
- 잠재 활동 h_t와 공간적 불확실성 맵을 통해 내부 다이나믹스를 시각화하여 의사결정 전략을 해석한다.
![Figure 1 : Computing a reaction time metric from a recurrent vision model. a. A schematic representation of training a cRNN with evidential deep learning (EDL; [ 30 ] ). Model outputs are interpreted as parameters ( $\boldsymbol{\alpha}$ ) of a Dirichlet distribution over class probability estimates](https://ar5iv.labs.arxiv.org/html/2306.11582/assets/x1.png)
실험 결과
연구 질문
- RQ1xi_cRNN가 인간에서 관찰되는 자극 의존적 반응 시간 패턴을 포착할 수 있는가?
- RQ2EDL로 학습된 안정적 cRNN이 다양한 시각 과제에서 인간 의사결정 시간과 정렬되는 시간적 다이내믹스를 보이는가?
- RQ3xi_cRNN가 과제 난이도, 자극 구조 또는 공간적 속성으로 인한 인간 RT 변동을 예측하는가?
주요 결과
- xi_cRNN는 자극 의존적 RT 패턴을 추적하고 네 가지 과제에서 인간 RT와 질적으로 정렬된다.
- xi_cRNN은 증가적 그룹화에서 인간 데이터와 유사한 순차적 주의집중 전략과 공간 비등방성을 드러낸다.
- xi_cRNN은 Planko 과제와 미로/경로 길이 조건에서 인간 RT와 상관되며 더 어려운 자극에 대해 더 긴 처리 시간을 반영한다.
- xi_cRNN은 장면 분류에서 구별가능성 관련 RT 경향을 예측하고 인간 RT와 상관관계가 있다(r = 0.19, p < .001).
- C-RBP를 이용한 안정적 학습은 시간에 적응적인 처리를 견고하게 만들어 BPTT 기반 접근법보다 일반화가 향상된다.
- 이 프레임워크는 모델 다이나믹스와 인간의 시간적 처리 간의 일반적인 비교 방법을 제공하고 신경과학 가설을 생성한다.
![Figure 2 : Human versus cRNN temporal alignment on an incremental grouping task. a. Description of the task (inspired by cognitive neuroscience studies [ 47 ] ). b. Visualization of the cRNN dynamics. The two lines represent the average latent trajectories across $1K$ validation stimuli labeled ‘‘ye](https://ar5iv.labs.arxiv.org/html/2306.11582/assets/figures/figure1b_lowres.png)
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.