[논문 리뷰] Analyzing Right-wing YouTube Channels: Hate, Violence and Discrimination
본 논문은 자막과 댓글을 기준 데이터셋과 비교하여 어휘적 분석, 주제 분석, 암시적 편향 분석을 통해 우익 YouTube 콘텐츠의 혐오, 폭력, 차별적 편향을 분석한다. 우익 채널에서 부정적인 언어의 비율이 높고 Muslims에 대한 특정 편향이 나타나며, 해설자(commentators)가 비디오 호스트(video hosts)보다 종종 더 Extreme한 경향이 있다.
As of 2018, YouTube, the major online video sharing website, hosts multiple channels promoting right-wing content. In this paper, we observe issues related to hate, violence and discriminatory bias in a dataset containing more than 7,000 videos and 17 million comments. We investigate similarities and differences between users' comments and video content in a selection of right-wing channels and compare it to a baseline set using a three-layered approach, in which we analyze (a) lexicon, (b) topics and (c) implicit biases present in the texts. Among other results, our analyses show that right-wing channels tend to (a) contain a higher degree of words from "negative" semantic fields, (b) raise more topics related to war and terrorism, and (c) demonstrate more discriminatory bias against Muslims (in videos) and towards LGBT people (in comments). Our findings shed light not only into the collective conduct of the YouTube community promoting and consuming right-wing content, but also into the general behavior of YouTube users.
연구 동기 및 목표
- Investigate hateful vocabulary, violent content, and discriminatory bias in a set of right-wing YouTube channels.
- Compare captions and comments within right-wing channels and against a baseline channel set.
- Propose a three-layered methodology (lexical, topic, implicit bias) using open-source tools for text analysis.
- Offer insights into both host and commenter behavior on YouTube regarding hate and discrimination.
제안 방법
- Three-layered approach: lexical analysis, topic analysis, and implicit bias analysis.
- Lexical: map words to Empath semantic fields; lemmatize; compute normalized category vectors per video and channel; measure caption-comment similarity via cosine similarity.
- Topic: apply Latent Dirichlet Allocation (LDA) to captions and comments to identify latent topics; use 300 topics with alpha=beta=1.0/num_topics.
- Implicit bias: construct Word Embedding Association Tests (WEAT) using word2vec embeddings trained on Wikipedia and domain data; compute effect sizes (Cohen’s d) and p-values via permutation tests.
- Data: collect 3,731 right-wing videos and 5,071,728 comments; baseline: 3,942 videos and 12,519,590 comments; English-filtered subsets used for analysis.
실험 결과
연구 질문
- RQ1RQ-1: Is hateful vocabulary, violent content, and discriminatory bias more accentuated in right-wing channels than baseline channels?
- RQ2RQ-2: Are commentators more or less exacerbated than video hosts in expressing hate and discrimination?
주요 결과
- Right-wing channels show higher fractions of negative semantic fields (e.g., aggression, kill, violence) than baseline channels.
- Topics in right-wing captions emphasize war/terrorism and information warfare; baseline topics are broader (celebrities, TV shows, etc.).
- Implicit bias analyses reveal stronger anti-Muslim bias in captions, and comparatively varied bias against LGBT people; baseline biases are amplified relative to Wikipedia in some cases.
- Commentators generally exhibit higher levels of hate-related language (disgust, hate, swearing) than video hosts in many cases.
- Similarity between captions and comments varies by channel; more popular channels tend to show higher caption-comment similarity.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.