Skip to main content
QUICK REVIEW

[论文解读] It's not what you said, it's how you said it: discriminative perception of speech as a multichannel communication system

Sarenne Wallbridge, Peter Bell|arXiv (Cornell University)|May 1, 2021
Speech and dialogue systems参考文献 39被引用 4
一句话总结

本研究探讨了听者如何利用词汇性与非词汇性语音通道来区分语调不同的、但词汇完全相同的语句。通过一项新颖的行为辨别任务,研究发现非词汇性(语调)语境显著提升了感知准确性,优于仅依赖词汇性语境的表现,表明语调在口语对话感知中起着关键作用。

ABSTRACT

People convey information extremely effectively through spoken interaction using multiple channels of information transmission: the lexical channel of what is said, and the non-lexical channel of how it is said. We propose studying human perception of spoken communication as a means to better understand how information is encoded across these channels, focusing on the question 'What characteristics of communicative context affect listener's expectations of speech?'. To investigate this, we present a novel behavioural task testing whether listeners can discriminate between the true utterance in a dialogue and utterances sampled from other contexts with the same lexical content. We characterize how perception - and subsequent discriminative capability - is affected by different degrees of additional contextual information across both the lexical and non-lexical channel of speech. Results demonstrate that people can effectively discriminate between different prosodic realisations, that non-lexical context is informative, and that this channel provides more salient information than the lexical channel, highlighting the importance of the non-lexical channel in spoken interaction.

研究动机与目标

  • 探究词汇性与非词汇性语音通道中的交际语境如何塑造口语对话中听者的预期。
  • 检验听者是否能够区分词汇相同但语调不同的语句。
  • 评估词汇性与非词汇性语境在支持语音辨别感知中的相对贡献。
  • 探究不同层次的语境信息(仅文本、音频+文本、未来语境)如何影响感知表现与信心水平。
  • 通过识别人类在实时对话感知中使用的显著线索,为自动语音表征的设计提供依据。

提出的方法

  • 设计一项新颖的行为辨别任务,让参与者判断目标语句是原始版本还是具有相同词汇内容的采样替代版本。
  • 使用Switchboard语料库中的刺激材料,从不同对话语境中选取词汇等价的语句。
  • 呈现不同语境模态(仅文本、音频+文本、未来语境)和语境范围(局部与扩展语境)的条件。
  • 测量在不同条件下辨别准确率与参与者信心水平,以评估对词汇性与非词汇性线索的依赖程度。
  • 分析表现趋势与信心评分,推断听者如何在不同语音通道间整合信息。
  • 控制说话人身份与词汇频率,以隔离语调与语境变化的影响。

实验结果

研究问题

  • RQ1非词汇性(语调)语境的增加如何影响听者对词汇等价语句的辨别能力?
  • RQ2仅依赖词汇性语境是否能提升辨别表现,还是语调语境更具信息量?
  • RQ3不同类型的语境信息(如前一话轮)与持续时间如何影响听者的预期?
  • RQ4参与者信心水平在不同语境条件下在多大程度上与实际表现一致?
  • RQ5该辨别任务能否揭示听者在实时对话感知中如何权衡词汇性与非词汇性线索的差异?

主要发现

  • 当提供非词汇性(语调)语境时,参与者辨别准确率显著提高,表明语调是语音感知中的显著线索。
  • 即使仅提供前一话轮的文本,表现依然很高,表明局部语调线索已足以实现辨别。
  • 随着语境信息的增加,参与者信心水平上升,但这种信心并不总与准确率正相关,尤其在仅提供文本的未来语境条件下。
  • 仅提供文本的未来语境条件(3,3)表现最低,接近随机基线,尽管信心水平很高,表明感知到的语境信息量与实际信息量存在错配。
  • 非词汇性语境提供的线索比仅依赖词汇性语境更具信息量,凸显其在感知辨别中的主导作用。
  • 结果表明,自监督语音表征学习方法可从引入局部语调语境作为训练信号中获益。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。