Skip to main content
QUICK REVIEW

[論文レビュー] Language Models Trained on Media Diets Can Predict Public Opinion

Eric Chu, Jacob Andreas|arXiv (Cornell University)|Mar 28, 2023
Computational and Text Analysis Methods被引用数 23
ひとこと要約

特定のメディアダイエットに適応した言語モデルは、COVID-19と消費者信頼に関するサブ集団の調査回答を予測でき、相関はおおよそ r=0.46、プロンプトやメディアソースに対する頑健性がある。

ABSTRACT

Public opinion reflects and shapes societal behavior, but the traditional survey-based tools to measure it are limited. We introduce a novel approach to probe media diet models -- language models adapted to online news, TV broadcast, or radio show content -- that can emulate the opinions of subpopulations that have consumed a set of media. To validate this method, we use as ground truth the opinions expressed in U.S. nationally representative surveys on COVID-19 and consumer confidence. Our studies indicate that this approach is (1) predictive of human judgements found in survey response distributions and robust to phrasing and channels of media exposure, (2) more accurate at modeling people who follow media more closely, and (3) aligned with literature on which types of opinions are affected by media consumption. Probing language models provides a powerful new method for investigating media effects, has practical applications in supplementing polls and forecasting public opinion, and suggests a need for further study of the surprising fidelity with which neural language models can predict human responses.

研究の動機と目的

  • サブ集団のメディアダイエットで訓練された言語モデルを用いて、世論を予測する新しい方法を提案する。
  • メディアコンテンツでファインチューニングしたモデルが、COVID-19と消費者信頼に関する調査回答の分布を予測できることを示す。
  • 質問の言い換えや異なるメディアソース、注目度の違いに対する頑健性を評価する。

提案手法

  • 特定のソース(オンラインニュース、テレビ、ラジオ)からのメディアダイエットデータセット上で、ベースとなる言語モデル(例:BERT)をファインチューニングする。
  • 調査質問に由来する空欄埋めプロンプトでメディアダイエットモデルを検証し、ターゲット単語の確率を計算する。
  • 同義語で確率を正規化・グループ化して、メディアダイエットスコアを導出する。
  • メディアダイエットスコア(および任意でニュースへの注目度)を実際の調査回答割合へ対応づける回帰を適合させる。
  • 埋め込み空間で最近傍分析を用いて、モデルの予測を訓練データに追跡する。

実験結果

リサーチクエスチョン

  • RQ1RQ1a: Do media-diet models have predictive power for survey responses?
  • RQ2RQ1b: Are pretrained neural models necessary, or are simpler models sufficient?
  • RQ3RQ1c: Is synonym-grouping necessary for scoring media diets?
  • RQ4RQ1d: Are results robust to paraphrasing of prompts?
  • RQ5RQ2: Do media-diet models vary with level of attention to news or across media sources?
  • RQ6RQ3: Are certain topics or question types more strongly predicted by media-diet models?

主な発見

  • Media-diet scores correlate with survey proportions (r = 0.458, 95% CI [0.350, 0.553]).
  • A regression using media-diet scores significantly predicts survey proportions (beta = 0.115, 95% CI [0.087, 0.142]).
  • Synonym-grouping improves correlations (r = 0.458 vs. r = 0.190 without grouping).
  • Models adapted to media content outperform baseline BERT (r = 0.274 for BERT alone; r = 0.458 for Media Diet BERT).
  • Combining media-diet scores with attention to news yields higher predictive power (beta = 0.523, CI [0.164, 0.882]; R2 = 0.3327).
  • Predictive power holds across online, TV, and radio sources, and across different prompts and paraphrase methods.

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。