Skip to main content
QUICK REVIEW

[论文解读] Annotation of Emotion Carriers in Personal Narratives

Aniruddha Tammewar, Alessandra Cervone|arXiv (Cornell University)|Feb 27, 2020
Sentiment Analysis and Opinion Mining参考文献 41被引用 8
一句话总结

本文提出了一种用于识别口语个人叙事中情感载体(即解释叙述者情感状态的文本片段)的新标注方案。基于德语USoMS语料库,四位标注员对转录的叙事文本中标注了情感载体,尽管存在主观性,但内容词的重叠度很高;分析显示填充词(如'ähm')经常出现在载体附近,表明其位置模式可为叙事理解系统中的自动情感载体检测提供帮助。

ABSTRACT

We are interested in the problem of understanding personal narratives (PN) - spoken or written - recollections of facts, events, and thoughts. In PN, emotion carriers are the speech or text segments that best explain the emotional state of the user. Such segments may include entities, verb or noun phrases. Advanced automatic understanding of PNs requires not only the prediction of the user emotional state but also to identify which events (e.g. "the loss of relative" or "the visit of grandpa") or people ( e.g. "the old group of high school mates") carry the emotion manifested during the personal recollection. This work proposes and evaluates an annotation model for identifying emotion carriers in spoken personal narratives. Compared to other text genres such as news and microblogs, spoken PNs are particularly challenging because a narrative is usually unstructured, involving multiple sub-events and characters as well as thoughts and associated emotions perceived by the narrator. In this work, we experiment with annotating emotion carriers from speech transcriptions in the Ulm State-of-Mind in Speech (USoMS) corpus, a dataset of German PNs. We believe this resource could be used for experiments in the automatic extraction of emotion carriers from PN, a task that could provide further advancements in narrative understanding.

研究动机与目标

  • 开发一种系统化的标注框架,用于识别口语个人叙事中的情感载体,这些叙事具有复杂性、非结构化和情感丰富等特点。
  • 通过识别触发或解释情感状态的具体文本片段,弥补情感分析中仅限于情感分类的空白。
  • 评估在具有挑战性且主观性强的任务中,对包含多个子事件和情感的长篇叙事文本进行标注的标注者间一致性和标注质量。
  • 探究语言线索(如不流畅现象,例如填充词)是否与情感载体存在系统性关联,从而可能有助于自动检测。
  • 创建一个用于训练和评估个人叙事中情感载体抽取的自动系统的研究资源,以支持更深层次的叙事理解。

提出的方法

  • 基于USoMS语料库中的口语个人叙事转录文本,由四位标注员手动标注情感载体,识别最能解释叙述者情感状态的文本片段。
  • 情感载体被定义为包含动词、名词或短语的多词跨度,用以传达叙述者情感的原因或触发因素(例如'Praktikum'或'positive Aufregung')。
  • 使用标准指标计算标注者间一致性,以评估可靠性,并通过重叠标注的定性分析评估一致性。
  • 开展位置分析,研究情感载体在叙事结构中的分布,重点关注平均位置和聚类模式。
  • 将填充词(如'ähm'、'äh'、'mhm')识别为不流畅现象,并分析其在每个载体±5个词符范围内的相对位置。
  • 通过对照实验,将填充词与情感载体的接近程度与随机内容词(名词、形容词、动词、副词)进行比较,以验证所观察到的位置模式的显著性。

实验结果

研究问题

  • RQ1在识别口语个人叙事中的情感载体时,标注者间的一致性如何?多位标注员的标注在多大程度上保持一致?
  • RQ2情感载体在叙事结构中是否均匀分布?还是倾向于集中在开头、中间或结尾等特定区域?
  • RQ3在口语叙事中,不流畅现象(如填充词)与情感载体之间是否存在统计学上显著的位置关联?
  • RQ4个人叙事中的情感载体是否倾向于为内容词(如名词、动词),而非功能词?这如何影响检测效果?
  • RQ5填充词与情感载体之间的接近程度是否可作为自动情感载体抽取在叙事理解系统中的可靠信号?

主要发现

  • 尽管任务具有主观性,标注者间的一致性仍达到中等水平,且多位标注员在内容词上重叠度很高,表明在识别有意义的情感载体方面具有一致性。
  • 情感载体最常出现在叙事的中段,其平均位置接近叙事中心,表明其在情感表达中具有核心作用。
  • 填充词如'ähm'和'äh'在情感载体后方第2位(+2)出现频率最高(占7%的案例),且在+1、+3以及-1、-4位置也表现出显著聚类。
  • 与随机内容词相比,填充词与情感载体的接近频率显著更高,表明二者之间存在非随机的位置关联。
  • 分析表明,情感载体通常由内容词构成,如'Praktikum'(实习)或'positive Aufregung'(积极兴奋),表明事件和情感短语是关键载体。
  • 填充词与情感载体之间观察到的模式支持了如下假设:不流畅现象可能标记认知或情感处理的时刻,或可为建模口语叙事中的情感意图提供帮助。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。