Skip to main content
QUICK REVIEW

[论文解读] Evaluation of short range depth sonifications for visual-to-auditory sensory substitution

Louis Commère, Jean Rouat|arXiv (Cornell University)|Apr 11, 2023
Tactile and Sensory Interactions被引用 4
一句话总结

本研究评估了五种深度声音化技术——频率、振幅、短促高音哨声的重复率、混响以及纯音与噪声的信噪比——在盲人用户中的视觉到听觉感官替代应用。通过使用视力正常但蒙眼的参与者,研究发现哨声重复率在深度感知的准确性与记忆性方面表现最佳,优于其他方法在定位精度和用户感知方面的表现。

ABSTRACT

Visual to auditory sensory substitution devices convert visual information into sound and can provide valuable assistance for blind people. Recent iterations of these devices rely on depth sensors. Rules for converting depth into sound (i.e. the sonifications) are often designed arbitrarily, with no strong evidence for choosing one over another. The purpose of this work is to compare and understand the effectiveness of five depth sonifications in order to assist the design process of future visual to auditory systems for blind people which rely on depth sensors. The frequency, amplitude and reverberation of the sound as well as the repetition rate of short high-pitched sounds and the signal-to-noise ratio of a mixture between pure sound and noise are studied. We conducted positioning experiments with twenty-eight sighted blindfolded participants. Stage 1 incorporates learning phases followed by depth estimation tasks. Stage 2 adds the additional challenge of azimuth estimation to the first stage's protocol. Stage 3 tests learning retention by incorporating a 10-minute break before re-testing depth estimation. The best depth estimates in stage 1 were obtained with the sound frequency and the repetition rate of beeps. In stage 2, the beep repetition rate yielded the best depth estimation and no significant difference was observed for the azimuth estimation. Results of stage 3 showed that the beep repetition rate was the easiest sonification to memorize. Based on statistical analysis of the results, we discuss the effectiveness of each sonification and compare with other studies that encode depth into sounds. Finally we provide recommendations for the design of depth encoding.

研究动机与目标

  • 评估并比较五种深度声音化策略在视觉到听觉感官替代系统中的有效性。
  • 确定哪种声音化方法能使盲人用户基于三维传感器数据实现最准确且最易记忆的深度估计。
  • 为依赖深度传感器的未来感官替代设备提供基于证据的设计建议。
  • 探究在学习、任务复杂度和记忆保持条件下,声音化性能是否存在差异。
  • 评估感知自然性与跨模态对应关系在声音化设计中的作用。

提出的方法

  • 开展三个实验阶段:(1) 带有学习过程的深度估计,(2) 深度与方位角联合估计,(3) 10分钟休息后的记忆保持测试。
  • 使用28名视力正常但蒙眼的参与者,以消除声音化任务中视觉反馈的影响。
  • 评估五种声音化方法:频率、振幅、短促高音哨声的重复率、混响,以及纯音与噪声混合的信噪比。
  • 测量绝对定位误差,并通过统计分析比较不同声音化方法的表现。
  • 计算偶然水平误差为(b−a)/3,以建立显著性检验的基线表现。
  • 收集关于每种声音化方法感知自然性与直观性的定性反馈。

实验结果

研究问题

  • RQ1在短距离(1米)环境中,哪种深度声音化方法能实现最准确的深度估计?
  • RQ2在包含方位角估计的情况下,不同声音化策略的表现如何变化?
  • RQ3在10分钟延迟后,哪种声音化方法最容易被记住,表明其长期可用性?
  • RQ4感知自然性与跨模态对应关系如何影响用户偏好与表现?
  • RQ5基于视力正常但蒙眼的参与者所得结果,能否推广至真实应用场景中的盲人用户?

主要发现

  • 在第一阶段中,哨声重复率产生了最准确的深度估计,其绝对定位误差显著低于其他声音化方法。
  • 在第二阶段(包含方位角估计)中,哨声重复率仍为表现最佳的声音化方法,且各方法在方位角估计准确性上无显著差异。
  • 哨声重复率被评价为编码深度最自然、最直观的方法,参与者一致偏好该方法。
  • 第三阶段的表现表明,哨声重复率是最易记忆的声音化方法,10分钟休息后误差最低。
  • 基于声音频率的声音化方法虽具备良好的定量准确性,但被认为不够直观,且与深度的自然关联性较弱。
  • 纯音与白噪声混合的声音化方法获得了强烈的定性反馈,显示出潜力,但其准确性仍低于哨声重复率。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。