[论文解读] The Follower Count Fallacy: Detecting Twitter Users with Manipulated Follower Count
本文提出一种无监督方法,通过基于时间特征和行为特征的局部邻域模型估算用户的真实关注者数量,以检测关注者数量被操纵的Twitter用户。该方法在关注者数量预测中达到84.2%的准确率,并通过测量预测值与实际值的偏差,以98.62%的精确度检测出被操纵的用户,表明其对黑市服务造成的合成性关注者膨胀具有高度容忍度。
Online Social Networks (OSN) are increasingly being used as platform for an effective communication, to engage with other users, and to create a social worth via number of likes, followers and shares. Such metrics and crowd-sourced ratings give the OSN user a sense of social reputation which she tries to maintain and boost to be more influential. Users artificially bolster their social reputation via black-market web services. In this work, we identify users which manipulate their projected follower count using an unsupervised local neighborhood detection method. We identify a neighborhood of the user based on a robust set of features which reflect user similarity in terms of the expected follower count. We show that follower count estimation using our method has 84.2% accuracy with a low error rate. In addition, we estimate the follower count of the user under suspicion by finding its neighborhood drawn from a large random sample of Twitter. We show that our method is highly tolerant to synthetic manipulation of followers. Using the deviation of predicted follower count from the displayed count, we are also able to detect customers with a high precision of 98.62%
研究动机与目标
- 检测使用黑市服务人为虚增关注者数量的Twitter用户。
- 基于用户的社会邻域,估算其真实且抗操纵的关注者数量。
- 评估预测方法对合成性关注者操纵的容忍度。
- 评估关注者操纵对Klout等影响力指标的影响,以及关注者数量估算的可靠性。
- 提供一种无需依赖真实关注者数量的鲁棒无监督检测框架。
提出的方法
- 该方法利用时间签名(如发帖频率和爆发模式)识别用户所在局部邻域,这些特征难以伪造。
- 通过从大量随机抽取的Twitter用户中聚类具有相似社会地位、时间行为和社交签名特征的用户,构建邻域。
- 采用关注者数量预测模型,基于用户邻域的关注者数量中位数或均值,估算其真实关注者数量。
- 将预测关注者数量与实际显示的关注者数量之间的偏差,用作检测操纵行为的信号。
- 通过在模拟账户上进行关注者操纵并测量预测关注者数量与Klout等影响力指标的变化,评估方法的容忍度。
- 定义容忍度指标为实际变化与预测变化的比值,0.0表示完全容忍,1.0表示无容忍。
实验结果
研究问题
- RQ1时间签名能否可靠地识别出具有非自然增长关注者数量的用户?
- RQ2基于邻域的关注者数量估算方法在多大程度上能抵抗黑市关注者服务的操纵?
- RQ3该方法在多大程度上能准确预测被操纵关注者数量的用户的真实关注者数量?
- RQ4在存在合成性关注者膨胀的情况下,预测关注者数量与Klout等既定影响力指标相比如何?
- RQ5关注者数量预测值与实际值的偏差能否作为可靠信号,用于检测购买关注者的用户?
主要发现
- 当允许预测值与显示值相差±100名关注者时,关注者数量预测方法的准确率达到84.2%。
- 该方法对操纵表现出高度容忍,真实客户的关注者数量预测值平均仅变化0.10,而Klout评分平均增加0.47。
- 对于通过免费增值型黑市服务创建的模拟账户,预测关注者数量平均仅变化0.011,而Klout评分平均增加0.32。
- 通过设定预测值与显示值之间10%的偏差阈值,该方法以98.62%的精确度检测出被操纵关注者数量的用户。
- 在操纵下,预测关注者数量保持稳定,而Klout等影响力指标对人工关注者膨胀高度敏感。
- 该方法在不同关注者数量群体(G1: ≤1000,G2: 1000–10,000,G3: ≥10,000)中均表现有效,展现出在不同用户规模下的稳定性能。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。