Skip to main content
QUICK REVIEW

[论文解读] Advancing an Interdisciplinary Science of Conversation: Insights from a Large Multimodal Corpus of Human Speech

Andrew Reece, Gus Cooney|arXiv (Cornell University)|Mar 1, 2022
Language, Metaphor, and Cognition被引用 4
一句话总结

本文通过分析包含1,656段录音英语对话、总计850小时及700多万词的大型多模态语料库,推动了对话的跨学科科学。研究提出了新型对话轮次分割方法,应用机器学习对语音、视觉和文本特征进行分析以预测对话成功与否,并揭示了对话动态与全生命周期内幸福感之间的关联。

ABSTRACT

People spend a substantial portion of their lives engaged in conversation, and yet our scientific understanding of conversation is still in its infancy. In this report we advance an interdisciplinary science of conversation, with findings from a large, novel, multimodal corpus of 1,656 recorded conversations in spoken English. This 7+ million word, 850 hour corpus totals over 1TB of audio, video, and transcripts, with moment-to-moment measures of vocal, facial, and semantic expression, along with an extensive survey of speaker post conversation reflections. We leverage the considerable scope of the corpus to (1) extend key findings from the literature, such as the cooperativeness of human turn-taking; (2) define novel algorithmic procedures for the segmentation of speech into conversational turns; (3) apply machine learning insights across various textual, auditory, and visual features to analyze what makes conversations succeed or fail; and (4) explore how conversations are related to well-being across the lifespan. We also report (5) a comprehensive mixed-method report, based on quantitative analysis and qualitative review of each recording, that showcases how individuals from diverse backgrounds alter their communication patterns and find ways to connect. We conclude with a discussion of how this large-scale public dataset may offer new directions for future research, especially across disciplinary boundaries, as scholars from a variety of fields appear increasingly interested in the study of conversation.

研究动机与目标

  • 通过整合语言学、计算机科学、心理学和社会学的洞见,建立对话的跨学科基础科学。
  • 解决尽管对话在日常生活中具有核心作用,但其科学理解仍有限的问题。
  • 开发并验证利用多模态数据将话语自动分割为对话轮次的新算法程序。
  • 分析多模态特征(语音、面部、语义)如何预测对话的成功或失败。
  • 探索不同年龄群体中对话行为与幸福感之间的纵向关系。

提出的方法

  • 收集大规模多模态语料库,包含音频、视频、转录的口语内容,以及对语音、面部和语义表达的逐时标注。
  • 应用机器学习模型,整合文本、音频和视觉特征,以预测对话结果。
  • 开发新型算法程序,利用多模态线索实现对话轮次的自动分割。
  • 整合对话后的调查数据,以丰富对说话者体验的定性和定量分析。
  • 采用混合方法分析,结合统计建模与个别录音的定性审查,以捕捉行为多样性。
  • 利用该语料库检验并拓展对话分析中的既定发现,例如对话轮次的协作性。

实验结果

研究问题

  • RQ1在自然的人际对话中,多模态线索(语音、面部、语义)如何共同变化?
  • RQ2在使用语音、视频和文本特征的情况下,机器学习模型在多大程度上能够预测对话成功?
  • RQ3来自不同背景的个体在对话模式上如何变化?他们使用何种策略来建立联系?
  • RQ4对话行为与不同年龄群体的幸福感之间存在何种关系?
  • RQ5如何利用多模态数据改进自动对话轮次分割?

主要发现

  • 语料证实了人类对话轮次的协作性,说话人转换与多模态线索之间存在强烈一致性。
  • 机器学习模型在结合语音、视觉和文本特征后,显著提升了识别成功对话的预测性能。
  • 来自不同背景的个体会调整其沟通模式——如语速、面部表情丰富度和轮次策略——以促进连接。
  • 在全生命周期中,对话参与度与自我报告的幸福感之间存在可测量的正相关关系。
  • 所提出的对话轮次分割算法方法通过整合实时语音和面部动态,优于传统单模态方法。
  • 对话后的反思揭示了参与者对自己对话贡献和情绪状态感知的一致性模式。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。