[Paper Review] ChatGPT-4 Outperforms Experts and Crowd Workers in Annotating Political Twitter Messages with Zero-Shot Learning
ChatGPT-4 is evaluated on classifying political party affiliation of Twitter posters and outperforms expert coders and MTurk crowd workers in accuracy and reliability, with comparable or lower bias, using zero-shot learning.
This paper assesses the accuracy, reliability and bias of the Large Language Model (LLM) ChatGPT-4 on the text analysis task of classifying the political affiliation of a Twitter poster based on the content of a tweet. The LLM is compared to manual annotation by both expert classifiers and crowd workers, generally considered the gold standard for such tasks. We use Twitter messages from United States politicians during the 2020 election, providing a ground truth against which to measure accuracy. The paper finds that ChatGPT-4 has achieves higher accuracy, higher reliability, and equal or lower bias than the human classifiers. The LLM is able to correctly annotate messages that require reasoning on the basis of contextual knowledge, and inferences around the author's intentions - traditionally seen as uniquely human abilities. These findings suggest that LLM will have substantial impact on the use of textual data in the social sciences, by enabling interpretive research at a scale.
Motivation & Objective
- Assess the accuracy of ChatGPT-4 for annotating political affiliation from Twitter content.
- Compare ChatGPT-4 performance to expert coders and MTurk crowd workers using ground-truth politician tweets.
- Evaluate reliability (inter-coder agreement) and bias in ChatGPT-4 versus human annotators.
- Use a ground-truth dataset of US senators' tweets from the 2020 election period as the evaluation baseline.
Proposed method
- Use 500 tweets (250 Republican, 250 Democrat) from US senators’ tweets before the 2020 US election, after filtering (no retweets/replies/URLs, length ≥100).
- Classify each tweet with ChatGPT-4 via API using zero-/few-shot prompting and run multiple times at different temperatures (5 runs at low temp 0.2 and 5 at high temp 1.0, totaling 5000 classifications).
- Compare ChatGPT-4 results to MTurk crowd workers (Master Qualified US workers, 10 annotators per tweet, with control questions) and to two expert classifiers.
- Compute accuracy against ground truth, Krippendorff’s Alpha for intercoder reliability, and bias (toward Democrat vs. Republican) across annotators.
Experimental results
Research questions
- RQ1Can ChatGPT-4 correctly infer political affiliation from tweet content using zero-shot learning?
- RQ2How does ChatGPT-4's accuracy compare to expert coders and MTurk crowd workers on this task?
- RQ3What is the reliability (intercoder agreement) of ChatGPT-4 relative to humans?
- RQ4Is there a bias toward predicting Democrat or Republican, and how does it compare across groups?
Key findings
- ChatGPT-4 achieves higher accuracy than both expert classifiers and MTurk workers.
- ChatGPT-4 shows higher intercoder reliability (Krippendorff’s Alpha) than human coders, especially at lower temperatures.
- All respondent groups (including ChatGPT-4 and experts) show a bias toward predicting Democrat, while MTurk workers exhibit a significantly stronger bias.
- ChatGPT-4's performance is demonstrated across tweets requiring contextual knowledge and inference of author intent.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.