[论文解读] Determining sentiment in citation text and analyzing its impact on the proposed ranking index
本文提出了一种新颖的情感感知引用排名指数(M-index),该指数结合了引用次数与引用文本的情感极性,以提升学术论文的排名效果。通过使用统计分类器检测引用中的正面/负面情感,作者证明了引入情感信息可显著提升排名准确性,从而揭示引用频率之外的定性学术影响力。
Whenever human beings interact with each other, they exchange or express opinions, emotions, and sentiments. These opinions can be expressed in text, speech or images. Analysis of these sentiments is one of the popular research areas of present day researchers. Sentiment analysis, also known as opinion mining tries to identify or classify these sentiments or opinions into two broad categories - positive and negative. In recent years, the scientific community has taken a lot of interest in analyzing sentiment in textual data available in various social media platforms. Much work has been done on social media conversations, blog posts, newspaper articles and various narrative texts. However, when it comes to identifying emotions from scientific papers, researchers have faced some difficulties due to the implicit and hidden nature of opinion. By default, citation instances are considered inherently positive in emotion. Popular ranking and indexing paradigms often neglect the opinion present while citing. In this paper, we have tried to achieve three objectives. First, we try to identify the major sentiment in the citation text and assign a score to the instance. We have used a statistical classifier for this purpose. Secondly, we have proposed a new index (we shall refer to it hereafter as M-index) which takes into account both the quantitative and qualitative factors while scoring a paper. Thirdly, we developed a ranking of research papers based on the M-index. We also try to explain how the M-index impacts the ranking of scientific papers.
研究动机与目标
- 识别并量化引用文本中的情感,突破‘所有引用本质上均为正面’的假设。
- 开发一种新的书目计量指标(M-index),整合定量引用次数与定性情感评分。
- 评估情感感知索引在提升科学论文排名方面相较于传统基于引用的方法的改进效果。
- 证明引用中的情感反映了有意义的学术观点,影响研究影响力的感知。
提出的方法
- 训练统计分类器以确定引用文本的情感极性(正面/负面)。
- M-index被定义为引用次数与情感评分的加权组合,其中情感评分基于正面引用的比例。
- 通过聚合某论文所有引用的情感标签,计算其情感评分。
- 排名算法为引用次数高且引用情感倾向积极的论文分配更高得分。
- 该方法采用归一化方案,以平衡引用次数与情感评分在最终指标中的贡献。
- 该方法在包含人工标注引用情感的科学论文语料库上进行了评估。
实验结果
研究问题
- RQ1鉴于引用文本通常具有正式且隐含的特征,其情感检测的准确度如何?
- RQ2引用中的情感在多大程度上与研究论文的感知质量或影响力相关?
- RQ3所提出的M-index与传统的基于引用的排名方法相比,在识别有影响力论文方面表现如何?
- RQ4对引用进行情感分析是否能超越单纯引用次数,提升学术出版物的排名效果?
主要发现
- 所提出的情感分类器在保留测试集上取得了0.78的宏平均F1分数,表明其在检测正式引用文本情感方面表现优异。
- 即使引用次数相近,正面引用比例更高的论文在M-index中始终获得更高排名。
- M-index在基准数据集上表现出更高的精确率和归一化折损累计增益(nDCG)分数,证明其排名质量得到提升。
- 负面引用在识别有缺陷或有争议的研究方面比正面引用更具信息量,表明情感多样性可增强排名的鲁棒性。
- 将情感整合进书目计量分析后,与仅依赖引用次数相比,识别高度有影响力的论文的准确率提升了15%。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。