Skip to main content
QUICK REVIEW

[论文解读] Over a Decade of Social Opinion Mining.

Keith Cortis, Brian Davis|arXiv (Cornell University)|Dec 5, 2020
Sentiment Analysis and Opinion Mining参考文献 405被引用 7
一句话总结

本篇系统性综述分析了2007至2018年间485项关于社交意见挖掘的研究,探讨了技术、平台、模态和语言,以从多模态社交媒体内容中提取情感、情绪、反讽及其他意见维度。研究识别出关键趋势、工具和研究空白,为推动人工智能在营销、政治、医疗保健和决策系统等领域的应用奠定基础。

ABSTRACT

Social media popularity and importance is on the increase, due to people using it for various types of social interaction across multiple channels. This social interaction by online users includes submission of feedback, opinions and recommendations about various individuals, entities, topics, and events. This systematic review focuses on the evolving research area of Social Opinion Mining, tasked with the identification of multiple opinion dimensions, such as subjectivity, sentiment polarity, emotion, affect, sarcasm and irony, from user-generated content represented across multiple social media platforms and in various media formats, like text, image, video and audio. Therefore, through Social Opinion Mining, natural language can be understood in terms of the different opinion dimensions, as expressed by humans. This contributes towards the evolution of Artificial Intelligence, which in turn helps the advancement of several real-world use cases, such as customer service and decision making. A thorough systematic review was carried out on Social Opinion Mining research which totals 485 studies and spans a period of twelve years between 2007 and 2018. The in-depth analysis focuses on the social media platforms, techniques, social datasets, language, modality, tools and technologies, natural language processing tasks and other aspects derived from the published studies. Such multi-source information fusion plays a fundamental role in mining of people's social opinions from social media platforms. These can be utilised in many application areas, ranging from marketing, advertising and sales for product/service management, and in multiple domains and industries, such as politics, technology, finance, healthcare, sports and government. Future research directions are presented, whereas further research and development has the potential of leaving a wider academic and societal impact.

研究动机与目标

  • 映射2007至2018年这12年间社交意见挖掘研究的演变历程,并识别关键研究趋势。
  • 分析意见挖掘研究中所用社交媒体平台、语言、模态(文本、图像、视频、音频)及数据集的多样性。
  • 评估用于挖掘多维意见(如情感极性、情绪、反讽)的技术、工具和自然语言处理任务。
  • 识别当前研究中的空白与未来研究方向,以推动社交意见挖掘及其现实世界应用的发展。
  • 通过整合多源信息融合在意见挖掘中的研究成果,支持更稳健的人工智系统开发。

提出的方法

  • 对2007至2018年间发表于同行评审期刊和会议的485项研究进行了系统性文献综述。
  • 根据社交媒体平台、语言、模态(文本、图像、视频、音频)以及情感分析、反讽检测等自然语言处理任务对研究进行分类。
  • 分析不同数据类型中意见挖掘所采用的技术,包括基于规则、机器学习和深度学习的方法。
  • 评估公开可用的社会数据集、工具和框架在意见挖掘研究中的使用情况,以评估可复现性与可扩展性。
  • 综合多个维度(平台、语言、模态和应用领域)的研究发现,以识别研究趋势与挑战。
  • 通过主题分析与出版趋势及方法论空白的定量分析,识别未来研究方向。

实验结果

研究问题

  • RQ12007至2018年间,社交意见挖掘研究中占主导地位的社交媒体平台和语言是什么?
  • RQ2在此期间,针对文本、图像、视频和音频模态,意见挖掘技术如何演变?
  • RQ3意见挖掘中最常见的自然语言处理任务和工具是什么?它们在不同平台和模态之间有何差异?
  • RQ4当前意见挖掘研究中的主要挑战与局限是什么,特别是针对反讽、讽刺和多模态融合?
  • RQ5哪些未来研究方向最有可能推动社交意见挖掘及其现实世界应用的发展?

主要发现

  • 大多数意见挖掘研究集中于Twitter和Facebook等平台的文本内容,对视频和音频模态的关注有限。
  • 英语是研究中的主导语言,非英语及低资源语言存在显著研究空白。
  • 机器学习与深度学习技术在情感和情绪检测任务中持续超越基于规则的方法,表现更优。
  • 反讽与讽刺检测仍是重大挑战,准确率较低,且可用的标注数据集有限。
  • 多模态意见挖掘(文本+图像/视频/音频)虽为新兴方向,但研究尚不充分,缺乏标准化数据集与工具。
  • 尽管在医疗保健、金融和政府等应用领域具有高潜在影响,但当前研究中这些领域仍代表性不足。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。