[论文解读] Visual Noise from Natural Scene Statistics Reveals Human Scene Category Representations
本文提出了REVEAL方法,这是一种新颖的技术,利用基于自然场景统计特性的视觉噪声和进化算法,可视化人类对真实世界场景类别(例如“街道”)的内部表征。通过让观察者将有噪声的刺激与他们的心理模板进行匹配,该方法重建出能够预测快速场景检测表现的主观场景表征,首次证明了复杂场景的私人心理图像可以被客观可视化和验证。
Our perceptions are guided both by the bottom-up information entering our eyes, as well as our top-down expectations of what we will see. Although bottom-up visual processing has been extensively studied, comparatively little is known about top-down signals. Here, we describe REVEAL (Representations Envisioned Via Evolutionary ALgorithm), a method for visualizing an observer's internal representation of a complex, real-world scene, allowing us to, for the first time, visualize the top-down information in an observer's mind. REVEAL rests on two innovations for solving this high dimensional problem: visual noise that samples from natural image statistics, and a computer algorithm that collaborates with human observers to efficiently obtain a solution. In this work, we visualize observers' internal representations of a visual scene category (street) using an experiment in which the observer views the naturalistic visual noise and collaborates with the algorithm to externalize his internal representation. As no scene information was presented, observers had to use their internal knowledge of the target, matching it with the visual features in the noise. We matched reconstructed images with images of real-world street scenes to enhance visualization. Critically, we show that the visualized mental images can be used to predict rapid scene detection performance, as each observer had faster and more accurate responses to detecting real-world images that were the most similar to his reconstructed street templates. These results show that it is possible to visualize previously unobservable mental representations of real world stimuli. More broadly, REVEAL provides a general method for objectively examining the content of previously private, subjective mental experiences.
研究动机与目标
- 开发一种可视化先前不可见的、主观的现实世界场景类别心理表征的方法。
- 研究基于对场景内部知识的自上而下的预期在缺乏清晰视觉输入时如何引导视觉知觉。
- 构建一个框架,通过人机协作实现对私人心理模板的客观、数据驱动的重建。
- 通过测试其对现实世界场景检测表现的预测能力,验证重建的心理表征。
- 证明复杂场景的心理表征可借助计算与感知反馈被外部化并进行定量分析。
提出的方法
- 该方法生成基于自然图像统计特性的视觉噪声,确保知觉上的真实感。
- 进化算法基于人类反馈,迭代优化有噪声的图像,引导观察者选择最接近其目标场景(例如“街道”)内部表征的噪声模式。
- 观察者查看不断演化的噪声图像,并选择最符合其内心模板的版本,提供持续反馈。
- 算法利用此反馈优化噪声模式,逐步收敛至反映观察者内部表征的重建图像。
- 最终的重建图像与真实世界场景图像进行匹配,以增强可解释性和验证性。
- 该方法利用了即使没有明确的场景内容,观察者也能基于其对场景结构和统计特性的内部知识匹配噪声模式的事实。
实验结果
研究问题
- RQ1我们能否客观地可视化观察者对复杂真实世界场景类别的内部主观表征?
- RQ2在缺乏清晰视觉输入的情况下,基于对场景先前知识的自上而下预期在多大程度上引导知觉?
- RQ3重建的心理表征能否预测快速场景检测任务中的表现?
- RQ4自然场景的统计特性在多大程度上影响人类场景表征的结构?
- RQ5无需直接场景线索,人机协作算法能否有效外部化私人心理模板?
主要发现
- REVEAL方法成功利用仅有的视觉噪声和人类反馈,重建了观察者对场景类别(例如“街道”)的特定个体心理表征。
- 重建的图像在知觉上连贯且与真实世界街道场景视觉相似,表明该方法捕捉到了有意义的内部表征。
- 观察者对与其自身重建模板最相似的真实世界街道图像表现出更快、更准确的检测表现。
- 真实场景与观察者重建模板之间的相似度程度,能高度准确地预测检测表现。
- 结果证实,对场景类别的自上而下知识显著影响视觉知觉,即使在缺乏清晰视觉特征的情况下亦然。
- 本研究证明,以往私密且主观的心理体验可借助基于自然场景统计的计算框架实现可视化和验证。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。