Skip to main content
QUICK REVIEW

[论文解读] ScenarioSA: A Large Scale Conversational Database for Interactive Sentiment Analysis

Yazhou Zhang, Lingling Song|arXiv (Cornell University)|Jul 12, 2019
Sentiment Analysis and Opinion Mining参考文献 6被引用 6
一句话总结

本文提出了ScenarioSA,一个大规模、公开可用的对话数据库,包含2,214个经过人工标注的多轮英语对话,用于交互式情感分析。该数据集捕捉了多样化现实场景中动态情感状态及人际情感演变,支持开发和评估超越静态文本分类的先进交互式情感分析模型。

ABSTRACT

Interactive sentiment analysis is an emerging, yet challenging, subtask of the sentiment analysis problem. It aims to discover the affective state and sentimental change of each person in a conversation. Existing sentiment analysis approaches are insufficient in modelling the interactions among people. However, the development of new approaches are critically limited by the lack of labelled interactive sentiment datasets. In this paper, we present a new conversational emotion database that we have created and made publically available, namely ScenarioSA. We manually label 2,214 multi-turn English conversations collected from natural contexts. In comparison with existing sentiment datasets, ScenarioSA (1) covers a wide range of scenarios; (2) describes the interactions between two speakers; and (3) reflects the sentimental evolution of each speaker over the course of a conversation. Finally, we evaluate various state-of-the-art algorithms on ScenarioSA, demonstrating the need of novel interactive sentiment analysis models and the potential of ScenarioSA to facilitate the development of such models.

研究动机与目标

  • 为解决多轮对话中交互式情感分析缺乏大规模标注数据集的问题。
  • 创建一个能够捕捉个体说话者在互动过程中动态情感状态和情感演变的数据集。
  • 支持开发能够建模人际情感动态而非孤立情感预测的新颖模型。
  • 为评估最先进算法在交互式情感理解任务中的表现提供基准。

提出的方法

  • 作者从真实情境中收集了2,214个自然发生的多轮英语对话,以确保多样性和真实性。
  • 每个对话均对每句发言进行了人工情感标签标注,以捕捉每位说话者随时间推移的情感状态。
  • 该数据集涵盖广泛的情境,包括社交、职场和情感互动,以反映现实生活对话的复杂性。
  • 明确标注了各轮之间的情感演变,以反映互动过程中情感状态的变化。
  • 数据集结构支持轮次级和对话级情感分析,能够建模说话者特定的情感轨迹。
  • 该数据集已公开发布,以支持交互式情感分析领域的可重现研究和模型开发。

实验结果

研究问题

  • RQ1现有最先进模型在捕捉多轮对话中交互式情感动态方面的表现如何?
  • RQ2现有情感分析模型能否有效建模个体说话者随时间推移的演变情感状态?
  • RQ3当应用于具有人际互动动态的对话情感分析时,现有方法存在哪些局限性?
  • RQ4数据集中情境多样性在多大程度上影响模型的泛化能力和性能?
  • RQ5大规模人工标注的对话数据集在多大程度上能提升交互式情感分析模型的开发?

主要发现

  • 在ScenarioSA上对最先进模型的评估揭示了显著的性能差距,表明当前模型在建模交互式情感动态方面能力不足。
  • 该数据集的复杂性和多样性暴露了现有方法将情感视为静态或说话者无关的局限性。
  • 结果明确表明需要开发专门设计用于捕捉对话中情感演变和人际互动的新颖模型。
  • ScenarioSA的可用性使得对交互式情感分析系统的评估更加准确和细致。
  • 该数据集支持开发能够追踪对话轮次中个体情感轨迹的模型,这是以往基准所缺乏的能力。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。