Skip to main content
QUICK REVIEW

[Paper Review] Creating Large Language Model Resistant Exams: Guidelines and Strategies

Simon kaare Larsen|arXiv (Cornell University)|Apr 18, 2023
Artificial Intelligence in Healthcare and EducationMedicine3 citations
TL;DR

This paper proposes practical guidelines for designing large language model (LLM)-resistant exams by leveraging real-world scenarios, deliberate inaccuracies, non-textual content, and soft skill assessment. It demonstrates that with strategic design, exams can resist LLM cheating while maintaining academic integrity and relevance to modern professional contexts.

ABSTRACT

The proliferation of Large Language Models (LLMs), such as ChatGPT, has raised concerns about their potential impact on academic integrity, prompting the need for LLM-resistant exam designs. This article investigates the performance of LLMs on exams and their implications for assessment, focusing on ChatGPT's abilities and limitations. We propose guidelines for creating LLM-resistant exams, including content moderation, deliberate inaccuracies, real-world scenarios beyond the model's knowledge base, effective distractor options, evaluating soft skills, and incorporating non-textual information. The article also highlights the significance of adapting assessments to modern tools and promoting essential skills development in students. By adopting these strategies, educators can maintain academic integrity while ensuring that assessments accurately reflect contemporary professional settings and address the challenges and opportunities posed by artificial intelligence in education.

Motivation & Objective

  • To address growing concerns about academic integrity due to the use of large language models (LLMs) like ChatGPT in student assessments.
  • To investigate the capabilities and limitations of LLMs in answering exam questions, particularly focusing on ChatGPT.
  • To develop actionable, evidence-based guidelines for educators to design exams that are resistant to LLM-generated answers.
  • To ensure assessments reflect real-world professional competencies and promote essential skills development in students.
  • To adapt educational assessment practices to the evolving landscape of AI tools in education.

Proposed method

  • Incorporate real-world scenarios that extend beyond the training data of LLMs, reducing their ability to generate plausible but incorrect answers.
  • Introduce deliberate inaccuracies into exam questions to test students' ability to detect and correct misinformation, a skill LLMs often fail to recognize.
  • Use non-textual information such as diagrams, charts, or images that LLMs cannot interpret accurately, increasing the difficulty of automated cheating.
  • Design questions that require evaluation of soft skills like critical thinking, ethical reasoning, and contextual judgment—areas where LLMs underperform.
  • Implement effective distractor options in multiple-choice questions to challenge LLMs' reasoning and reduce the likelihood of random correct answers.
  • Structure assessments to emphasize application over recall, favoring tasks that require personal reflection or situational adaptation.

Experimental results

Research questions

  • RQ1How do large language models like ChatGPT perform on traditional exam formats, and where do they exhibit vulnerabilities?
  • RQ2What specific design features in exams can reduce the likelihood of successful LLM-generated cheating?
  • RQ3In what ways can assessments be restructured to prioritize human skills that LLMs cannot replicate?
  • RQ4How can educators incorporate real-world, context-specific scenarios to limit LLM effectiveness?
  • RQ5What role do non-textual elements and deliberate inaccuracies play in increasing exam resistance to LLMs?

Key findings

  • LLMs such as ChatGPT can generate plausible but incorrect answers on standard exam questions, especially when the content is ambiguous or lacks real-world context.
  • The inclusion of deliberate inaccuracies in questions effectively exposes LLMs' inability to detect and correct false information, reducing their success rate in answering correctly.
  • Exams incorporating non-textual content like diagrams or charts significantly reduce the risk of LLM-based cheating, as LLMs cannot reliably interpret visual data.
  • Assessments emphasizing soft skills such as ethical judgment and critical thinking show lower susceptibility to LLM-generated responses, as these require human-level reasoning.
  • Real-world, scenario-based questions that require situational adaptation are less likely to be answered correctly by LLMs due to their reliance on training data rather than lived experience.
  • The proposed guidelines collectively enhance exam integrity by shifting assessment focus from recall to application, judgment, and contextual understanding.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.