[论文解读] The Role of User Profile for Fake News Detection
本文通過分析社交媒體上的分享行為,研究用戶檔案在假新聞檢測中的作用,以識別更可能分享假新聞或真實新聞的用戶。研究提取了顯式和隱式用戶檔案特徵(如發文數、政治傾向和人格特質),並證明這些特徵顯著提升了假新聞分類性能,在多種學習演算法下平均F1得分超過0.90,其中隱式特徵比顯式特徵更具區分性。
Consuming news from social media is becoming increasingly popular. Social media appeals to users due to its fast dissemination of information, low cost, and easy access. However, social media also enables the widespread of fake news. Because of the detrimental societal effects of fake news, detecting fake news has attracted increasing attention. However, the detection performance only using news contents is generally not satisfactory as fake news is written to mimic true news. Thus, there is a need for an in-depth understanding on the relationship between user profiles on social media and fake news. In this paper, we study the challenging problem of understanding and exploiting user profiles on social media for fake news detection. In an attempt to understand connections between user profiles and fake news, first, we measure users' sharing behaviors on social media and group representative users who are more likely to share fake and real news; then, we perform a comparative analysis of explicit and implicit profile features between these user groups, which reveals their potential to help differentiate fake news from real news. To exploit user profile features, we demonstrate the usefulness of these user profile features in a fake news classification task. We further validate the effectiveness of these features through feature importance analysis. The findings of this work lay the foundation for deeper exploration of user profile features of social media and enhance the capabilities for fake news detection.
研究动机与目标
- 理解社交媒體用戶檔案與分享假新聞或真實新聞之間的關係。
- 識別並描述區分分享假新聞與分享真實新聞的用戶的檔案特徵(顯式與隱式)。
- 評估用戶檔案特徵在提升假新聞檢測性能方面相較於基於內容的方法的有效性。
- 提供一種系統化且可解釋的框架,用於在假新聞檢測中利用用戶檔案,從而增強模型的魯棒性與解釋能力。
提出的方法
- 作者從Politifact和GossipCop構建了兩個真實世界數據集,並使用可靠的真實標籤來標註假新聞與真實新聞。
- 利用絕對與相對指標衡量用戶的分享行為,以識別更可能分享假新聞或真實新聞的代表性用戶群組。
- 顯式檔案特徵(如發文數、粉絲數)直接從用戶元數據中提取。
- 隱式檔案特徵(如政治傾向、人格特質、年齡、性別)則透過自然語言處理與機器學習技術,基於用戶生成的文本與社交信號推斷得出。
- 進行比較性統計分析,以評估分享假新聞與真實新聞的用戶在特徵分佈上的差異。
- 透過多種學習演算法的分類實驗驗證用戶檔案特徵的實用性,並進行特徵重要性分析以評估其穩健性與貢獻度。
实验结果
研究问题
- RQ1哪些用戶更可能分享假新聞或真實新聞?
- RQ2更可能分享假新聞或真實新聞的用戶具有哪些特徵,其檔案特徵是否存在顯著差異?
- RQ3用戶檔案特徵能否有效用於假新聞檢測?與基於內容的特徵相比表現如何?
主要发现
- 分享假新聞的用戶與分享真實新聞的用戶在檔案特徵上存在顯著差異,顯式與隱式特徵均具統計顯著性。
- 隱式檔案特徵(如政治傾向與人格特質)在區分假新聞與真實新聞方面表現出比顯式特徵更高的區分能力。
- 用戶檔案特徵在多種學習演算法下持續提升假新聞檢測性能,平均F1得分超過0.90。
- 所提出的用戶檔案特徵在假新聞分類任務中優於當前先進的基於內容的特徵。
- 特徵重要性分析確認,隱式特徵對檢測的貢獻度顯著高於顯式特徵,凸顯其在模型可解釋性與魯棒性方面的價值。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。