Skip to main content
QUICK REVIEW

[Paper Review] ChatGPT: The End of Online Exam Integrity?

Teo Sušnjak|arXiv (Cornell University)|Dec 19, 2022
Artificial Intelligence in Healthcare and EducationMedicine357 citations
TL;DR

The paper analyzes ChatGPT’s ability to perform high-level cognitive tasks and generate human-like text, and discusses its implications for online exam integrity and potential mitigations.

ABSTRACT

This study evaluated the ability of ChatGPT, a recently developed artificial intelligence (AI) agent, to perform high-level cognitive tasks and produce text that is indistinguishable from human-generated text. This capacity raises concerns about the potential use of ChatGPT as a tool for academic misconduct in online exams. The study found that ChatGPT is capable of exhibiting critical thinking skills and generating highly realistic text with minimal input, making it a potential threat to the integrity of online exams, particularly in tertiary education settings where such exams are becoming more prevalent. Returning to invigilated and oral exams could form part of the solution, while using advanced proctoring techniques and AI-text output detectors may be effective in addressing this issue, they are not likely to be foolproof solutions. Further research is needed to fully understand the implications of large language models like ChatGPT and to devise strategies for combating the risk of cheating using these tools. It is crucial for educators and institutions to be aware of the possibility of ChatGPT being used for cheating and to investigate measures to address it in order to maintain the fairness and validity of online exams for all students.

Motivation & Objective

  • Assess ChatGPT’s ability to generate and answer high-quality undergraduate-level critical thinking questions across disciplines.
  • Evaluate how ChatGPT can critically evaluate its own responses using universal intellectual standards.
  • Explore implications for the integrity of online exams in higher education and assess current mitigation strategies.
  • Suggest directions for future research and policy to address risks to assessment fairness.

Proposed method

  • Create a ChatGPT account and prompt it to generate difficult critical-thinking questions for multiple disciplines.
  • Have ChatGPT provide detailed answers to its own questions and then critically evaluate those answers.
  • Apply universal intellectual standards (relevance, clarity, accuracy, precision, depth, breadth, logic, persuasiveness, originality) to assess the responses.
  • Analyze responses across Education, Machine Learning, History, and Marketing to illustrate capabilities and limitations.
  • Discuss implications for online exams and the effectiveness of proctoring and AI-detection tools as mitigations.

Experimental results

Research questions

  • RQ1Can ChatGPT generate challenging, discipline-specific critical-thinking questions for undergraduates?
  • RQ2Can ChatGPT produce coherent, well-structured answers to its own questions?
  • RQ3Can ChatGPT critically evaluate its own responses and provide constructive improvement suggestions?
  • RQ4What are the implications of ChatGPT’s capabilities for online exam integrity and current mitigation strategies?

Key findings

  • ChatGPT can generate discipline-specific, challenging critical-thinking questions for undergraduates.
  • ChatGPT can produce detailed, coherent 500-word answers to its own questions.
  • ChatGPT can critically evaluate its own answers, listing strengths, weaknesses, and suggestions for improvement.
  • The capabilities demonstrated pose a potential threat to online exam integrity in higher education.
  • Current mitigation strategies (invigilated exams, proctoring, AI-detection) may not be foolproof against AI-generated cheating; further research is needed.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.