Skip to main content
QUICK REVIEW

[Paper Review] Have We Reached AGI? Comparing ChatGPT, Claude, and Gemini to Human Literacy and Education Benchmarks

Mfon Akpan|arXiv (Cornell University)|Jul 11, 2024
Artificial Intelligence in Healthcare and EducationMedicine3 citations
TL;DR

This study evaluates whether large language models (LLMs) like ChatGPT, Claude, and Gemini have approached Artificial General Intelligence (AGI) by benchmarking their performance against U.S. human literacy and education standards. Using data from the U.S. Census Bureau and technical reports, it finds that LLMs significantly outperform average Americans in undergraduate knowledge and reading comprehension tasks, suggesting substantial progress toward AGI, though broader cognitive assessments remain necessary for definitive conclusions.

ABSTRACT

Recent advancements in AI, particularly in large language models (LLMs) like ChatGPT, Claude, and Gemini, have prompted questions about their proximity to Artificial General Intelligence (AGI). This study compares LLM performance on educational benchmarks with Americans' average educational attainment and literacy levels, using data from the U.S. Census Bureau and technical reports. Results show that LLMs significantly outperform human benchmarks in tasks such as undergraduate knowledge and advanced reading comprehension, indicating substantial progress toward AGI. However, true AGI requires broader cognitive assessments. The study highlights the implications for AI development, education, and societal impact, emphasizing the need for ongoing research and ethical considerations.

Motivation & Objective

  • To assess whether recent large language models (LLMs) like ChatGPT, Claude, and Gemini have approached Artificial General Intelligence (AGI) by comparing their performance to human educational benchmarks.
  • To evaluate LLM capabilities against standardized measures of American literacy and educational attainment using data from the U.S. Census Bureau and technical reports.
  • To determine the extent to which LLMs surpass human averages in knowledge and comprehension tasks relevant to higher education.
  • To identify gaps in current benchmarks and highlight the need for broader cognitive assessments to define true AGI.
  • To inform AI development, education policy, and ethical considerations by quantifying LLM performance relative to human performance.

Proposed method

  • The study uses publicly available performance data from LLMs (ChatGPT, Claude, Gemini) on standardized educational and literacy benchmarks.
  • Human performance benchmarks are drawn from U.S. Census Bureau data on average educational attainment and literacy levels in the United States.
  • Comparative analysis is conducted between LLM outputs and human averages across domains such as undergraduate-level knowledge and advanced reading comprehension.
  • The evaluation focuses on tasks requiring general knowledge, inference, and comprehension, reflecting core components of human cognitive ability.
  • Performance is quantified using standardized metrics from educational assessments, enabling direct comparison with national human benchmarks.
  • The study emphasizes the limitations of narrow benchmarks and calls for broader, more comprehensive cognitive evaluations to assess AGI readiness.

Experimental results

Research questions

  • RQ1To what extent do current LLMs like ChatGPT, Claude, and Gemini outperform the average American in educational and literacy benchmarks?
  • RQ2How do LLMs compare to human performance in tasks requiring advanced reading comprehension and general knowledge?
  • RQ3What does the performance gap between LLMs and humans suggest about the current trajectory toward Artificial General Intelligence (AGI)?
  • RQ4Which cognitive domains are currently well-met by LLMs, and which remain underdeveloped relative to human capabilities?
  • RQ5What are the implications of LLM performance exceeding human averages for AI development, education, and societal ethics?

Key findings

  • LLMs such as ChatGPT, Claude, and Gemini significantly outperform the average American in undergraduate-level knowledge tasks, indicating strong performance on general education benchmarks.
  • In advanced reading comprehension tasks, LLMs demonstrate performance levels that exceed the average educational attainment of U.S. adults.
  • The study finds that LLMs achieve high scores on standardized assessments that measure general cognitive ability and factual knowledge, surpassing human averages in these domains.
  • Despite strong performance in specific domains, the study cautions that current benchmarks do not fully capture the breadth of human cognition required for true AGI.
  • The results suggest that while LLMs are approaching or exceeding human-level performance in narrow educational tasks, broader cognitive assessments are still needed to validate AGI claims.
  • The findings highlight the urgency of developing more comprehensive evaluation frameworks to assess whether AGI has truly been reached.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.