[Paper Review] TGDataset: Collecting and Exploring the Largest Telegram Channels Dataset
This paper introduces TGDataset, the largest publicly available collection of Telegram channels to date, comprising 120,979 channels and over 400 million messages. Using a snowball sampling method starting from 180 seed channels, the authors collected and analyzed multilingual content, revealing widespread dissemination of conspiracy theories, extremist ideologies, and borderline content, offering a critical resource for studying misinformation and platform dynamics on Telegram.
Telegram is one of the most popular instant messaging apps in today's digital age. In addition to providing a private messaging service, Telegram, with its channels, represents a valid medium for rapidly broadcasting content to a large audience (COVID-19 announcements), but, unfortunately, also for disseminating radical ideologies and coordinating attacks (Capitol Hill riot). This paper presents the TGDataset, a new dataset that includes 120,979 Telegram channels and over 400 million messages, making it the largest collection of Telegram channels to the best of our knowledge. After a brief introduction to the data collection process, we analyze the languages spoken within our dataset and the topic covered by English channels. Finally, we discuss some use cases in which our dataset can be extremely useful to understand better the Telegram ecosystem, as well as to study the diffusion of questionable news. In addition to the raw dataset, we released the scripts we used to analyze the dataset and the list of channels belonging to the network of a new conspiracy theory called Sabmyk.
Motivation & Objective
- To address the lack of large-scale, publicly available datasets on Telegram channels for academic research.
- To collect and characterize a comprehensive dataset of public Telegram channels across multiple languages and topics.
- To investigate the prevalence of controversial content, including conspiracy theories, extremist ideologies, and illegal activities, on Telegram.
- To provide researchers with a reproducible, open-source dataset and analysis tools to study misinformation and platform evolution.
- To enable longitudinal and cross-linguistic studies of Telegram's role in content dissemination and community formation.
Proposed method
- Employed a snowball sampling technique starting from 180 seed Telegram channels to iteratively discover linked channels via message forwarding.
- Collected public channel metadata including title, username, description, subscriber count, creation date, and message content.
- Applied automated language detection to identify the dominant languages used across channels in the dataset.
- Conducted topic modeling and manual labeling on English-language channels to identify prevalent themes and communities.
- Released the raw dataset, analysis scripts, and labeled data (language and topic) as open-source resources.
- Excluded images, links, and user-specific data to comply with ethical standards and avoid copyrighted or adult content.
Experimental results
Research questions
- RQ1What is the linguistic diversity of public Telegram channels, and which languages dominate the platform?
- RQ2What are the primary topics and communities present in English-language Telegram channels?
- RQ3How has the growth and political leaning of Telegram channels evolved over time, particularly following major platform policy changes?
- RQ4To what extent do Telegram channels host content related to conspiracy theories, extremist ideologies, or illegal activities?
- RQ5How can the TGDataset support research on misinformation, coordinated disinformation campaigns, and platform-based radicalization?
Key findings
- The TGDataset contains 120,979 public Telegram channels and over 400 million messages, making it the largest known public collection of its kind.
- Russian channels were the most active in the early years, but English channels overtook them in 2021, likely due to user migration from WhatsApp after its privacy policy change.
- Language detection revealed significant multilingual diversity, with English, Russian, and Arabic being among the most prevalent languages.
- Topic analysis of English channels uncovered a substantial presence of channels promoting white supremacy, revenge porn, carding, hacking, and conspiracy theories such as Sabmyk.
- The dataset includes channels engaged in illegal or harmful activities, including those coordinating market manipulation (e.g., pump-and-dump schemes) and spreading extremist content.
- The dataset's release enables future research on platform evolution, content moderation, and the spread of questionable information on privacy-focused platforms like Telegram.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.