Skip to main content
QUICK REVIEW

[论文解读] Assessment of SE-specific Sentiment Analysis Tools: An Extended Replication Study.

Nicole Novielli, Fabio Calefato|arXiv (Cornell University)|Oct 20, 2020
Hate Speech and Cyberbullying Detection被引用 4
一句话总结

本研究通过复制并扩展关于GitHub和Stack Overflow讨论中情感分析的先前研究,评估了软件工程(SE)专用的情感分析工具。研究发现,现成的SE工具在细粒度层面经常产生矛盾的结果,表明在SE实证研究中得出有效结论,必须针对特定平台进行调优或重新训练。

ABSTRACT

Sentiment analysis methods have become popular for investigating human communication, including discussions related to software projects. Since general-purpose sentiment analysis tools do not fit well with the information exchanged by software developers, new tools, specific for software engineering (SE), have been developed. We investigate to what extent SE-specific tools for sentiment analysis mitigate the threats to conclusion validity of empirical studies in software engineering, highlighted by previous research. First, we replicate two studies addressing the role of sentiment in security discussions on GitHub and in question-writing on Stack Overflow. Then, we extend the previous studies by assessing to what extent the tools agree with each other and with the manual annotation on a gold standard of 600 documents. We find that different SE-specific sentiment analysis tools might lead to contradictory results at a fine-grain level, when used 'off-the-shelf'. Conversely, platform-specific tuning or retraining might be needed to take into account differences in platform conventions, jargon, or document lengths.

研究动机与目标

  • 评估SE专用情感分析工具是否能提高实证SE研究中结论的有效性。
  • 复制先前关于GitHub安全讨论中情感分析和Stack Overflow问题撰写中情感分析的两项研究。
  • 评估多种工具在600篇文档的金标准数据集上与人工标注的一致性。
  • 调查现成的SE工具在不同SE平台(如GitHub和Stack Overflow)上是否能产生可靠结果。

提出的方法

  • 复制先前两项研究,聚焦于GitHub安全讨论中的情感分析和Stack Overflow问题撰写中的情感分析。
  • 创建一个由600篇文档组成的手动标注情感的金标准数据集。
  • 将多种SE专用情感分析工具应用于同一数据集,以比较其输出结果。
  • 使用标注者间一致性度量指标,对工具的一致性进行定量评估,并与人工标注结果进行比较。
  • 分析差异以识别原因,如平台惯例、领域术语或文档长度差异。
  • 评估平台特定调优或微调对工具性能和一致性的影响。

实验结果

研究问题

  • RQ1当应用于SE专用文本时,SE专用情感分析工具之间的一致性程度如何?
  • RQ2SE专用工具在金标准数据集上与人工情感标注的对齐程度如何?
  • RQ3当在不同SE平台(如GitHub和Stack Overflow)上直接使用时,SE工具是否能产生一致的结果?
  • RQ4导致SE工具在情感分类中出现差异的因素有哪些?
  • RQ5是否需要对平台进行特定调优或微调,才能在SE环境中实现可靠的情感分析?

主要发现

  • 当直接应用于细粒度SE文本时,不同的SE专用情感分析工具经常产生矛盾的情感分类结果。
  • 工具间的一致性较低,表明即使在相同SE文本上,工具之间也存在显著不一致。
  • SE工具与人工标注的对齐程度有限,表明在实证SE研究中可能存在可靠性问题。
  • 差异部分源于平台惯例、领域特定术语以及不同平台间文档长度的差异。
  • 本研究证明,必须对平台进行特定调优或微调,才能提高工具的一致性和有效性。
  • 由于语境和语言差异,现成的SE工具无法在不针对特定SE平台进行适应的情况下被可靠使用。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。