[论文解读] ANGUS: Real-time manipulation of vocal roughness for emotional speech transformations
ANGUS 是一种实时、计算高效的算法,通过振幅调制和时域滤波来操控语音和非语音声音的嗓音粗糙度,以模拟嗓音唤醒度。该方法能有效增强语音、动物叫声及乐器声音的感知情绪负面性,且听者无法可靠区分处理后与原始声音。
Vocal arousal, the non-linear acoustic features taken on by human and animal vocalizations when highly aroused, has an important communicative function because it signals aversive states such as fear, pain or distress. In this work, we present a computationally-efficient, real-time voice transformation algorithm, ANGUS, which uses amplitude modulation and time-domain filtering to simulate roughness, an important component of vocal arousal, in arbitrary voice recordings. In a series of 4 studies, we show that ANGUS allows parametric control over the spectral features of roughness like the presence of sub-harmonics and noise; that ANGUS increases the emotional negativity perceived by listeners, to a comparable level as a non-real-time analysis/resynthesis algorithm from the state-of-the-art; that listeners cannot distinguish transformed and non-transformed sounds above chance level; and that ANGUS has a similar emotional effect on animal vocalizations and musical instrument sounds than on human vocalizations. A real-time implementation of ANGUS is made available as open-source software, for use in experimental emotion reseach and affective computing.
研究动机与目标
- 开发一种实时、计算高效的嗓音粗糙度操控方法,用于情绪化语音转换。
- 解决在预录语音中添加非线性嗓音特征(如亚谐波和噪声)的实时解决方案缺失问题。
- 实现对任意音频信号中粗糙度成分(如亚谐波和频谱噪声)的参数化控制。
- 评估 ANGUS 对人类、动物及乐器声音情绪影响的效果,并与非实时的最先进方法进行比较。
- 评估听者对处理后与原始音频之间自然度和可区分性的感知。
提出的方法
- 使用低频调制信号对原始信号进行振幅调制,生成亚谐波分量。
- 使用时域滤波器整形调制信号的频谱包络,突出粗糙度特征。
- 将载波信号建模为余弦波,调制信号建模为深度 h ∈ [0,1] 的余弦波,产生位于 (ωc ± ωm) 处的边带。
- 输出信号被分解为原始载波和两个边带:y(t) = x_c(t) + y^+(t) + y^-(t),从而产生亚谐波能量。
- 将该算法集成至实时处理流水线中,实现对输入音频流的动态操控。
- 将 ANGUS 与最先进方法中的非实时分析-重合成方法进行比较,评估其在情绪转换保真度方面的表现。
实验结果
研究问题
- RQ1ANGUS 在多大程度上能够对与嗓音粗糙度相关的频谱特征(如亚谐波和噪声)实现参数化控制?
- RQ2ANGUS 是否能将感知情绪负面性提升至与非实时最先进方法相当的水平?
- RQ3在感知测试中,听者能否可靠地区分 ANGUS 处理后的语音与未处理语音?
- RQ4ANGUS 对非语音声音(如动物叫声、乐器)是否产生与人类语音相似的情绪效应?
- RQ5语境和声学因素(如音高、音量)如何与 ANGUS 产生的粗糙度在情绪感知中相互作用?
主要发现
- ANGUS 通过振幅调制和滤波成功实现了对粗糙度特征(如亚谐波和噪声)的参数化控制。
- 听者感知 ANGUS 处理后的语音情绪负面性显著增强,达到与非实时最先进分析-重合成方法相当的水平。
- 听者无法在高于随机水平的程度上区分 ANGUS 处理后的语音与未处理语音,表明其具有高度的感知自然度。
- ANGUS 在人类语音、动物叫声及乐器声音上的情绪效应具有一致性,证明其具有跨领域可迁移性。
- 平均音高对感知负面性的影响为 Δ = +1.0,表明音高与粗糙度可能作为独立但可叠加的情绪线索起作用。
- 作为基线的 CONTROL 算法被评价为比未处理刺激更自然,可能由于感知偏差或在短时长(600ms)剪辑中声学复杂度降低。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。