Skip to main content
QUICK REVIEW

[Paper Review] ChatGPT: A Meta-Analysis after 2.5 Months

Christoph Leiter, Ran Zhang|arXiv (Cornell University)|Feb 20, 2023
Artificial Intelligence in Healthcare and Education23 citations
TL;DR

The paper analyzes over 300k tweets and more than 150 scientific papers to assess ChatGPT’s public perception, sentiment trajectories, and research themes 2.5 months after its release, finding overall high perceived quality with language and topic-based variations and a mix of opportunities and threats in academia.

ABSTRACT

ChatGPT, a chatbot developed by OpenAI, has gained widespread popularity and media attention since its release in November 2022. However, little hard evidence is available regarding its perception in various sources. In this paper, we analyze over 300,000 tweets and more than 150 scientific papers to investigate how ChatGPT is perceived and discussed. Our findings show that ChatGPT is generally viewed as of high quality, with positive sentiment and emotions of joy dominating in social media. Its perception has slightly decreased since its debut, however, with joy decreasing and (negative) surprise on the rise, and it is perceived more negatively in languages other than English. In recent scientific papers, ChatGPT is characterized as a great opportunity across various fields including the medical domain, but also as a threat concerning ethics and receives mixed assessments for education. Our comprehensive meta-analysis of ChatGPT's current perception after 2.5 months since its release can contribute to shaping the public debate and informing its future development. We make our data available.

Motivation & Objective

  • Assess how ChatGPT is perceived across social media and scientific literature.
  • Quantify sentiment, emotion, and topic distributions over time after release.
  • Identify language-based and topic-based differences in perception.
  • Characterize how researchers frame ChatGPT as opportunity or threat across domains.
  • Provide data and annotations to inform public debate and development directions.

Proposed method

  • Collect over 334,808 tweets using the #ChatGPT hashtag and deduplicate robot accounts.
  • Translate non-English tweets to English using a Facebook multilingual model.
  • Classify tweet sentiment with a multilingual XLM-Roberta model trained on 198 million tweets (F1=71% for English).
  • Infer weekly sentiment, language-specific trends, and topic distributions using an English topic classifier trained on 124 million tweets with 19 classes.
  • Annotate a sample of tweets for sentiment and emotion using GoEmotions-based classifiers and manual checks.
  • Analyze Arxiv and SemanticScholar papers (≈150 papers) via abstract-based annotation across quality, topic, and social impact.
Figure 1: Upper: weekly average of sentiment overall language (solid line), over English tweets (dotted line) and non-English tweets (dashed line). Lower: Tweet counts distribution and sentiment percentage change at weekly level aggregation.
Figure 1: Upper: weekly average of sentiment overall language (solid line), over English tweets (dotted line) and non-English tweets (dashed line). Lower: Tweet counts distribution and sentiment percentage change at weekly level aggregation.

Experimental results

Research questions

  • RQ1What is the overall sentiment toward ChatGPT on social media and how does it evolve over the first 2.5 months after release?
  • RQ2How do sentiment and topics differ across languages and over time?
  • RQ3What themes (science & technology, education, news, diaries, business) dominate discussions and how do they relate to sentiment?
  • RQ4How do scientific papers describe ChatGPT in terms of quality, topic, and social impact?
  • RQ5What are the notable limitations and strengths highlighted by different sources across domains?

Key findings

  • Social media sentiment shows an overall downward trend after an initial rise, with English tweets more positive than non-English tweets.
  • Positive sentiment peaks early and declines slightly; neutral sentiment increases over time.
  • Joy and surprise emotions dominate non-neutral tweets, with joy decreasing over time and surprise generally increasing after updates.
  • English-language tweets show more positive sentiment than German, French, Spanish, and Japanese, with topic distributions explaining some language differences.
  • Arxiv and SemanticScholar papers largely rate ChatGPT as high-quality (4-5) and view it as an opportunity in several domains, though ethics and education are more contested as threats or mixed impacts.
  • Analyses indicate education-related papers express both opportunity and threat concerns, while ethics papers skew toward threat; overall, research attention is rising across both Arxiv and SemanticScholar.
Figure 2: Weekly sentiment distribution averaged per language
Figure 2: Weekly sentiment distribution averaged per language

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.