Skip to main content
QUICK REVIEW

[论文解读] Uncovering the Dark Side of Telegram: Fakes, Clones, Scams, and Conspiracy Movements

Massimo La Morgia, Alessandro Mei|arXiv (Cornell University)|Nov 26, 2021
Spam and Phishing Detection参考文献 27被引用 10
一句话总结

本文提出 TGDataset,一个包含 35,382 个 Telegram 频道的大规模快照,并研究了利用隐私功能进行伪造、克隆和诈骗的频道。该研究提出了一种机器学习模型,可实现 86% 的准确率来检测伪造频道,并分析了通过克隆和伪造频道传播的 Sabmyk 共识理论,揭示了广泛存在的边缘活动,包括信用卡盗用、儿童色情内容和加密货币诈骗。

ABSTRACT

Telegram is one of the most used instant messaging apps worldwide. Some of its success lies in providing high privacy protection and social network features like the channels -- virtual rooms in which only the admins can post and broadcast messages to all its subscribers. However, these same features contributed to the emergence of borderline activities and, as is common with Online Social Networks, the heavy presence of fake accounts. Telegram started to address these issues by introducing the verified and scam marks for the channels. Unfortunately, the problem is far from being solved. In this work, we perform a large-scale analysis of Telegram by collecting 35,382 different channels and over 130,000,000 messages. We study the channels that Telegram marks as verified or scam, highlighting analogies and differences. Then, we move to the unmarked channels. Here, we find some of the infamous activities also present on privacy-preserving services of the Dark Web, such as carding, sharing of illegal adult and copyright protected content. In addition, we identify and analyze two other types of channels: the clones and the fakes. Clones are channels that publish the exact content of another channel to gain subscribers and promote services. Instead, fakes are channels that attempt to impersonate celebrities or well-known services. Fakes are hard to identify even by the most advanced users. To detect the fake channels automatically, we propose a machine learning model that is able to identify them with an accuracy of 86%. Lastly, we study Sabmyk, a conspiracy theory that exploited fakes and clones to spread quickly on the platform reaching over 1,000,000 users.

研究动机与目标

  • 揭示 Telegram 上伪造、克隆和诈骗频道的普遍性,这些频道利用平台的隐私和匿名功能。
  • 分析公共频道中诸如信用卡盗用、非法成人内容和白人至上主义内容等边缘活动。
  • 开发一种高准确率的自动化检测伪造频道的机器学习模型。
  • 以 Sabmyk 共识理论为案例研究,探究利用伪造和克隆频道进行协调性虚假信息传播的机制。
  • 发布 TGDataset,这是首个通用、公开可用的 Telegram 频道数据集,供未来研究使用。

提出的方法

  • 通过公共 API 访问和网络爬取,收集了 35,382 个 Telegram 频道及超过 1.3 亿条消息。
  • 基于消息转发模式构建网络图,以追踪频道之间的信息传播路径。
  • 提出一种机器学习模型,利用文本和结构特征(如用户名相似度、消息模式)对伪造频道进行分类。
  • 在经验证和报告的诈骗频道标注数据上训练模型,实现 86% 的准确率。
  • 通过图分析方法识别 Sabmyk 网络结构,包括 98 个频道及其传播路径。
  • 对克隆和伪造频道进行内容分析,以识别其盈利模式和意识形态推广策略。

实验结果

研究问题

  • RQ1Telegram 频道中普遍存在哪些类型的边缘活动和非法活动?它们如何利用平台功能?
  • RQ2伪造和克隆频道在 Telegram 上的结构、内容和传播机制上存在哪些差异?
  • RQ3机器学习在多大程度上能够准确检测伪造的 Telegram 频道?所提出的模型效果如何?
  • RQ4Sabmyk 共识理论如何通过伪造和克隆频道迅速传播?其传播的网络模式是什么?
  • RQ5未经验证的频道对用户隐私、虚假信息传播和平台治理有何影响?

主要发现

  • 研究识别出 83 个克隆频道,这些频道通过复制官方频道的内容来吸引订阅者,并推广加密货币或意识形态。
  • 机器学习模型在检测伪造频道方面实现了 86% 的准确率,识别出 191 个可疑伪造频道,其中 27 个被确认为欺诈行为。
  • 诈骗频道可通过少于 2 跳的转发路径从经验证频道到达,使用户暴露于高风险内容。
  • Sabmyk 共识理论网络由 98 个频道组成,其中 98% 的成员通过与主频道的 1 跳连接加入。
  • 超过 100 万名用户被 Sabmyk 网络触及,该网络利用伪造和克隆频道迅速放大其信息传播。
  • 该数据集揭示了广泛存在的非法内容,包括儿童色情、报复性色情、信用卡盗用和白人至上主义材料,尽管 Telegram 拥有验证和诈骗标记系统。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。