Skip to main content
QUICK REVIEW

[論文レビュー] Computing a human-like reaction time metric from stable recurrent vision models

Lore Goetschalckx, Lakshmi Narasimhan Govindarajan|arXiv (Cornell University)|Jun 20, 2023
Visual perception and processing mechanisms参考文献 65被引用数 8
ひとこと要約

論文は、安定な再帰視覚モデルを用いた evidential deep learning で訓練された刺激計算可能な反応時間代理指標 xi_cRNN を導出し、人間の RT パターンと四つの視覚タスクにおける整合性を示す。

ABSTRACT

The meteoric rise in the adoption of deep neural networks as computational models of vision has inspired efforts to "align" these models with humans. One dimension of interest for alignment includes behavioral choices, but moving beyond characterizing choice patterns to capturing temporal aspects of visual decision-making has been challenging. Here, we sketch a general-purpose methodology to construct computational accounts of reaction times from a stimulus-computable, task-optimized model. Specifically, we introduce a novel metric leveraging insights from subjective logic theory summarizing evidence accumulation in recurrent vision models. We demonstrate that our metric aligns with patterns of human reaction times for stimulus manipulations across four disparate visual decision-making tasks spanning perceptual grouping, mental simulation, and scene categorization. This work paves the way for exploring the temporal alignment of model and human visual strategies in the context of various other cognitive tasks toward generating testable hypotheses for neuroscience. Links to the code and data can be found on the project page: https://serre-lab.github.io/rnn_rts_site.

研究の動機と目的

  • ニューラルネットワークのダイナミクスと人間の視覚的意思決定の時間的一致性を動機づける。
  • cRNN からモデル駆動の刺激計算可能な反応時間指標を開発する。
  • 反応時間指標が複数タスクで人間の RT パターンと定性的に一致することを示す。
  • 時間ダイナミクスを研究し神経科学的仮説を生成する枠組みを提供する。

提案手法

  • C-RBP と Evidential Deep Learning(EDL)を用いて安定な再帰視覚モデル(cRNN)を訓練し、クラスに対するディリクレ分布の信念を取得する。
  • RT 指標 xi_cRNN を、時間に対するモデル不確実性曲線の面積として定義する。xi_cRNN = integral_0^T Epsilon(t) dt。
  • アトラクタダイナミクスと EDL を用いて、追加の監視なしに時間発展する不確実性を得る。
  • 人間の RT との整合性を評価するために、4 つのタスクへこの枠組みを適用する:逐次的グルーピング、視覚シミュレーション(Planko)、迷路経路推論、シーンカテゴリー化。
  • 潜在的活動 h_t や空間的不確実性マップを可視化し、意思決定戦略を解釈する。
Figure 1 : Computing a reaction time metric from a recurrent vision model. a. A schematic representation of training a cRNN with evidential deep learning (EDL; [ 30 ] ). Model outputs are interpreted as parameters ( $\boldsymbol{\alpha}$ ) of a Dirichlet distribution over class probability estimates
Figure 1 : Computing a reaction time metric from a recurrent vision model. a. A schematic representation of training a cRNN with evidential deep learning (EDL; [ 30 ] ). Model outputs are interpreted as parameters ( $\boldsymbol{\alpha}$ ) of a Dirichlet distribution over class probability estimates

実験結果

リサーチクエスチョン

  • RQ1xi_cRNN は humans が観察する刺激依存の反応時間パターンを捉えられるか?
  • RQ2EDL で訓練された安定な cRNN は、人間の意思決定時間と整合する時間発展を示すか?
  • RQ3xi_cRNN は課題の難易度、刺激構造、空間的性質による人間の RT の変動を予測できるか?

主な発見

  • xi_cRNN は刺激依存の RT パターンを追跡し、4 つのタスクで人間の RT と定性的に一致する。
  • xi_cRNN はIncremental Grouping において連続的な注意のような戦略と、人間データに類似した空間的異方性を示す。
  • xi_cRNN は Planko タスクと迷路/経路長条件で人間の RT と相関し、より難しい刺激で処理時間が長くなることを反映する。
  • xi_cRNN はシーンカテゴリー化の識別性に関連する RT 傾向を予測し、人間の RT と相関(r = 0.19, p < .001)。
  • C-RBP を用いた安定訓練は堅牢で時間適応的な処理を生み出し、BPTT ベースのアプローチより一般化を向上させる。
  • この枠組みは、モデルのダイナミクスと人間の時間的処理を比較し神経科学的仮説を生成する汎用的手法を提供する。
Figure 2 : Human versus cRNN temporal alignment on an incremental grouping task. a. Description of the task (inspired by cognitive neuroscience studies [ 47 ] ). b. Visualization of the cRNN dynamics. The two lines represent the average latent trajectories across $1K$ validation stimuli labeled ‘‘ye
Figure 2 : Human versus cRNN temporal alignment on an incremental grouping task. a. Description of the task (inspired by cognitive neuroscience studies [ 47 ] ). b. Visualization of the cRNN dynamics. The two lines represent the average latent trajectories across $1K$ validation stimuli labeled ‘‘ye

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。