Skip to main content
QUICK REVIEW

[论文解读] Twitter and Polls: Analyzing and estimating political orientation of Twitter users in India General #Elections2014

Abhishek Bhola|arXiv (Cornell University)|Jun 19, 2014
Complex Network Analysis Techniques参考文献 27被引用 9
一句话总结

本研究分析了1821万条印度2014年大选期间的推文,采用四种方法对用户的政治倾向进行分类:基于内容的分类(文本和话题标签)、基于用户特征的分类,以及基于转发和提及网络的网络社区检测。基于网络的方法准确率超过80%,显著优于基于内容的方法,表明社交网络结构是比推文内容本身更可靠的政党倾向指标。

ABSTRACT

This year (2014) in the month of May, the tenure of the 15th Lok Sabha was to end and the elections to the 543 parliamentary seats were to be held. A whooping $5 billion were spent on these elections, which made us stand second only to the US Presidential elections in terms of money spent. Swelling number of Internet users and Online Social Media (OSM) users could effect 3-4% of urban population votes as per a report of IAMAI (Internet & Mobile Association of India). Our count of tweets related to elections from September 2013 to May 2014, was close to 18.21 million. We analyzed the complete dataset and found that the activity on Twitter peaked during important events. It was evident from our data that the political behavior of the politicians affected their followers count. Yet another aim of our work was to find an efficient way to classify the political orientation of the users on Twitter. We used four different techniques: two were based on the content of the tweets, one on the user based features and another based on community detection algorithm on the retweet and user mention networks. We found that the community detection algorithm worked best. We built a portal to show the analysis of the tweets of the last 24 hours. To the best of our knowledge, this is the first academic pursuit to analyze the elections data and classify the users in the India General Elections 2014.

研究动机与目标

  • 分析印度2014年大选期间推文的体量、时间分布及情感倾向。
  • 识别大选期间推特上政治话语和用户参与度的模式。
  • 开发并评估多种分类技术,以估算用户的政治倾向。
  • 构建一个实时门户网站,用于监控与选举相关的推特数据和情感倾向。

提出的方法

  • 使用推特的流媒体API,从2013年9月至2014年5月收集了1821万条与选举相关的推文。
  • 对内容分类应用TF-IDF向量化和基于话题标签的特征。
  • 使用用户层面的特征,如粉丝数、好友数以及与政党相关关键词的出现频率进行分类。
  • 构建转发网络和用户提及网络,应用社区检测算法对政治倾向进行分类。
  • 使用人工标注的1000个用户资料数据集(613个支持者,425个反对者)评估分类性能。
  • 开发了一个实时网络门户,用于可视化热门话题、情感趋势、地理位置以及政治网络动态。

实验结果

研究问题

  • RQ1印度2014年大选期间,推文活动量与重大选举事件之间是否存在相关性?
  • RQ2基于内容的特征(文本和话题标签)能否有效将用户分类为支持或反对BJP/国大党?
  • RQ3用户层面的行为特征(粉丝数、好友数、提及频率)在多大程度上可预测其政治倾向?
  • RQ4在转发和提及图上应用基于网络的社区检测方法,在识别政治倾向方面效果如何?
  • RQ5对推特数据进行实时分析,能否为选举期间的政治情感和受欢迎程度提供可操作的见解?

主要发现

  • 推文活动量在关键选举事件期间达到高峰,其中周二和周三的活跃度最高,尤其集中在下午晚些时候。
  • 基于TF-IDF和话题标签的内容分类方法,对支持者和反对者的分类准确率分别仅为42.42%和37.25%。
  • 基于用户特征的分类方法将支持者的准确率提升至50%以上,反对者的准确率提升至40%以上。
  • 在转发和提及网络上应用基于网络的社区检测方法,准确率超过80%,显著优于基于内容的方法。
  • 研究证实,将反对政治倾向进行分类比支持倾向更具挑战性,表明数据或语言使用中存在偏差。
  • 实时门户成功可视化了热门话题、情感趋势以及与选举相关的推文地理分布。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。