[論文レビュー] Measuring and modeling the perception of natural and unconstrained gaze in humans and machines
本論文は、現実世界の自然で制約のない状況における人間と機械の視線認識を調査し、動的ヒントがなくても、対面での相互作用において人間が優れていることを示している。目の領域からの入力のみを学習することで、深層学習モデルが人間の知覚パターン——例えばウォルスタントン錯視への感受性——を再現している。
Humans are remarkably adept at interpreting the gaze direction of other individuals in their surroundings. This skill is at the core of the ability to engage in joint visual attention, which is essential for establishing social interactions. How accurate are humans in determining the gaze direction of others in lifelike scenes, when they can move their heads and eyes freely, and what are the sources of information for the underlying perceptual processes? These questions pose a challenge from both empirical and computational perspectives, due to the complexity of the visual input in real-life situations. Here we measure empirically human accuracy in perceiving the gaze direction of others in lifelike scenes, and study computationally the sources of information and representations underlying this cognitive capacity. We show that humans perform better in face-to-face conditions compared with recorded conditions, and that this advantage is not due to the availability of input dynamics. We further show that humans are still performing well when only the eyes-region is visible, rather than the whole face. We develop a computational model, which replicates the pattern of human performance, including the finding that the eyes-region contains on its own, the required information for estimating both head orientation and direction of gaze. Consistent with neurophysiological findings on task-specific face regions in the brain, the learned computational representations reproduce perceptual effects such as the Wollaston illusion, when trained to estimate direction of gaze, but not when trained to recognize objects or faces.
研究の動機と目的
- 生々しく制約のない視覚的状況における人間の視線方向認識の正確さを実証的に測定すること。
- 人間の視線認識の背後にある視覚的情報源(例:頭部の向きや目の領域)を特定すること。
- 視線推定タスクにおいて人間の性能を再現する計算モデルを開発すること。
- 脳内のタスク固有の神経表現が、学習された人工的表現に類似しているかどうかを調査すること。
- 動的入力(例:頭部の動き)が人間の視線認識優位性に寄与しているかどうかを特定すること。
提案手法
- 対面と録画映像の条件を比較する人間の心理物理学実験を実施し、視線認識の正確さを測定した。
- 目の領域のみを用いて人間のパフォーマンスを収集し、その寄与を分離した。
- 階層的視覚処理を模倣したアーキテクチャを用いて、顔画像から視線方向を推定する深層畳み込みニューラルネットワーク(CNN)を訓練した。
- 顔認識や物体認識などの代替タスクで同じモデルを訓練し、学習された表現を比較した。
- ウォルスタントン錯視のような知覚的錯視に対するモデルのパフォーマンスを評価し、人間の知覚と一致するかを検証した。
- 活性化解析を用いて、モデルの内部表現が顔処理領域に関する神経生理学的発見をどれだけ模倣しているかを検証した。
実験結果
リサーチクエスチョン
- RQ1リアルタイムの対面相互作用と録画映像刺激との間で、人間の視線認識の正確さに差は生じるか?
- RQ2自然な状況下での正確な視線推定に最も寄与する視覚的手がかり(例:頭部の向きや目の領域)は何か?
- RQ3視線推定タスクに特化して訓練された深層学習モデルは、人間の知覚的パターン、特に視覚的錯視への感受性を再現できるか?
- RQ4視線推定のために学習された表現は、顔認識や物体認識のために学習された表現と異なるか?また、それらはタスク固有の神経組織に類似しているか?
- RQ5人間の視線認識における優位性は、入力の動的要因に起因するのか、それとも文脈的・社会的ヒントなどの他の要因に起因するのか?
主な発見
- 動的ヒントを除去しても、対面での人間の視線認識の正確さは録画条件よりも顕著に優れている。
- 目の領域のみが視線推定に十分な情報を含んでおり、この領域のみが可視である場合でも、人間のパフォーマンスは高いままである。
- 視線方向推定に特化して訓練された深層学習モデルは、人間のパフォーマンスパターン——ウォルスタントン錯視への感受性を含めて——を再現している。
- モデルの内部表現は神経生理学的発見と類似している:視線推定のためのタスク固有の特徴が発達するが、顔認識や物体認識のタスクで訓練した場合にはそうならない。
- 人間の視線認識における優位性は、入力の動的要因に起因しないことが示唆され、より高次の文脈的・社会的処理に依存している可能性がある。
- モデルが知覚的錯視を再現できたという成功は、その内部表現が人間の視線認識のための視覚処理の主要な側面を捉えていることを示している。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。