[论文解读] Texture Interpolation for Probing Visual Perception
本文提出了一种基于最优传输理论的新型纹理插值方法,应用于深度卷积神经网络(CNN)激活,表明纹理分布可被椭圆分布良好近似。该方法实现了纹理之间的感知平滑插值,人类与猕猴神经数据证实,V4神经元线性编码插值参数,而V1则不能,揭示了纹理感知几何的神经基础。
Texture synthesis models are important tools for understanding visual processing. In particular, statistical approaches based on neurally relevant features have been instrumental in understanding aspects of visual perception and of neural coding. New deep learning-based approaches further improve the quality of synthetic textures. Yet, it is still unclear why deep texture synthesis performs so well, and applications of this new framework to probe visual perception are scarce. Here, we show that distributions of deep convolutional neural network (CNN) activations of a texture are well described by elliptical distributions and therefore, following optimal transport theory, constraining their mean and covariance is sufficient to generate new texture samples. Then, we propose the natural geodesics (ie the shortest path between two points) arising with the optimal transport metric to interpolate between arbitrary textures. Compared to other CNN-based approaches, our interpolation method appears to match more closely the geometry of texture perception, and our mathematical framework is better suited to study its statistical nature. We apply our method by measuring the perceptual scale associated to the interpolation parameter in human observers, and the neural sensitivity of different areas of visual cortex in macaque monkeys.
研究动机与目标
- 开发一种基于深度神经网络激活的数学上严谨的纹理插值方法。
- 研究视觉皮层中纹理感知与神经编码如何与纹理空间的几何结构相关。
- 检验基于最优传输推导的插值路径是否比现有方法更符合感知与神经反应。
- 使用MLDS协议测量人类受试者对纹理插值的感知尺度。
- 比较猕猴视觉皮层区域(V1与V4)对插值参数的神经敏感性。
提出的方法
- 将纹理分布建模为CNN激活统计量(均值与协方差)空间中的椭圆分布。
- 应用最优传输理论,在这些分布的空间上定义黎曼度量,从而实现在纹理之间的测地线插值。
- 使用激活分布之间的Wasserstein距离作为纹理生成与插值的损失函数。
- 在CNN特征的均值与协方差向量空间中进行线性插值,以生成中间纹理。
- 使用线性判别分析(LDA)从神经动作电位发放序列与发放次数中解码插值参数。
- 使用MLDS(最大似然差异标度)协议测量人类观察者对插值的完整感知尺度。
实验结果
研究问题
- RQ1基于最优传输在纹理之间进行插值是否能产生与人类感知一致的感知平滑过渡?
- RQ2V1与V4中的神经元对插值参数的敏感性是否不同,这种敏感性是否反映了对纹理空间几何的神经编码?
- RQ3由CNN激活统计量定义的纹理空间几何是否能预测感知与神经反应?
- RQ4训练过的与未训练的CNN如何影响插值路径的平滑性与感知质量?
- RQ5与Portilla-Simoncelli方法中使用的异质性汇总统计量相比,使用均值与协方差的同质特征空间是否能更好地建模纹理感知?
主要发现
- 纹理的CNN激活分布可被椭圆分布良好近似,支持使用均值与协方差作为充分统计量。
- 在CNN激活统计量空间中通过最优传输进行插值,可在纹理之间产生感知平滑的过渡。
- 人类受试者在MLDS协议测量下表现出插值的线性感知尺度,表明对纹理空间的一致感知采样。
- V4中的神经元使用发放次数与完整发放序列均表现出对插值参数的强线性解码性能,表明其对纹理过渡的动态编码。
- V1中的神经元对插值参数无显著解码性能,表明其对中间纹理状态不敏感。
- 训练过的CNN产生的插值路径比未训练的更平滑,且这些路径由稳定、同质的纹理组成,表明对流形的近似能力得到提升。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。