Skip to main content
QUICK REVIEW

[论文解读] ChatGPT: A Meta-Analysis after 2.5 Months

Christoph Leiter, Ran Zhang|arXiv (Cornell University)|Feb 20, 2023
Artificial Intelligence in Healthcare and Education被引用 23
一句话总结

本文在发布后2.5个月分析了超过 300k 条推文和超过 150 篇科学论文,以评估 ChatGPT 的公众认知、情感轨迹和研究主题,结果显示总体感知质量较高,但存在语言与主题差异,以及学术界对机会与威胁的混合态势。

ABSTRACT

ChatGPT, a chatbot developed by OpenAI, has gained widespread popularity and media attention since its release in November 2022. However, little hard evidence is available regarding its perception in various sources. In this paper, we analyze over 300,000 tweets and more than 150 scientific papers to investigate how ChatGPT is perceived and discussed. Our findings show that ChatGPT is generally viewed as of high quality, with positive sentiment and emotions of joy dominating in social media. Its perception has slightly decreased since its debut, however, with joy decreasing and (negative) surprise on the rise, and it is perceived more negatively in languages other than English. In recent scientific papers, ChatGPT is characterized as a great opportunity across various fields including the medical domain, but also as a threat concerning ethics and receives mixed assessments for education. Our comprehensive meta-analysis of ChatGPT's current perception after 2.5 months since its release can contribute to shaping the public debate and informing its future development. We make our data available.

研究动机与目标

  • 评估 ChatGPT 在社交媒体和科学文献中的认知态度。
  • 量化发布后随时间变化的情感、情绪与主题分布。
  • 识别基于语言和主题的感知差异。
  • 描述研究人员在各领域如何将 ChatGPT 框定为机遇或威胁。
  • 提供数据和注释以促进公共讨论和发展方向的决策。

提出的方法

  • 使用 #ChatGPT 标签收集超过 334,808 条推文,并对机器人账户进行去重。
  • 使用 Facebook 的多语言模型将非英语推文翻译为英语。
  • 使用在 198 百万条推文上训练的多语言 XLM-Roberta 模型对推文情感进行分类(英语 F1=71%)。
  • 使用在 124 百万条推文、19 个类别上训练的英语主题分类器推断每周情感、语言特异性趋势和主题分布。
  • 使用基于 GoEmotions 的分类器和人工检查对推文样本进行情感与情绪注释。
  • 通过摘要注释分析 Arxiv 和 SemanticScholar 的论文(≈150 篇),覆盖质量、主题和社会影响。
Figure 1: Upper: weekly average of sentiment overall language (solid line), over English tweets (dotted line) and non-English tweets (dashed line). Lower: Tweet counts distribution and sentiment percentage change at weekly level aggregation.
Figure 1: Upper: weekly average of sentiment overall language (solid line), over English tweets (dotted line) and non-English tweets (dashed line). Lower: Tweet counts distribution and sentiment percentage change at weekly level aggregation.

实验结果

研究问题

  • RQ1在社交媒体上对 ChatGPT 的总体情感如何,以及在发布后的前 2.5 个月内如何演变?
  • RQ2情感和主题如何在语言之间以及随时间变化?
  • RQ3哪些主题(科学与技术、教育、新闻、日记、商业)主导讨论,以及它们与情感之间的关系?
  • RQ4科学论文如何在质量、主题和社会影响方面描述 ChatGPT?
  • RQ5跨领域的不同来源突出的显著限制和优势是什么?

主要发现

  • 社交媒体情感在初步上升后总体呈下降趋势,英语推文比非英语推文更积极。
  • 积极情感峰值较早出现并略有下降;中性情感随时间增加。
  • 快乐与惊讶情绪在非中性推文中占主导,快乐随时间下降,惊讶在更新后通常上升。
  • 英语推文的情感在总体上比德语、法语、西班牙语和日语更积极,且主题分布解释了一些语言差异。
  • Arxiv 与 SemanticScholar 的论文大多将 ChatGPT 评为高质量(4-5),并在若干领域视其为机会,尽管伦理与教育方面的影响更具争议,属于威胁或混合影响。
  • 分析表明教育相关论文同时表达机会与威胁担忧,而伦理学论文倾向于威胁;总体而言,Arxiv 和 SemanticScholar 的研究关注度在上升。
Figure 2: Weekly sentiment distribution averaged per language
Figure 2: Weekly sentiment distribution averaged per language

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。