[Paper Review] Game of Tones: Faculty detection of GPT-4 generated content in university assessments
The study tests GPT-4 content in university assessments and evaluates faculty ability to detect it with Turnitin AI detection, revealing gaps in detection and suggesting assessment reforms.
This study explores the robustness of university assessments against the use of Open AI's Generative Pre-Trained Transformer 4 (GPT-4) generated content and evaluates the ability of academic staff to detect its use when supported by the Turnitin Artificial Intelligence (AI) detection tool. The research involved twenty-two GPT-4 generated submissions being created and included in the assessment process to be marked by fifteen different faculty members. The study reveals that although the detection tool identified 91% of the experimental submissions as containing some AI-generated content, the total detected content was only 54.8%. This suggests that the use of adversarial techniques regarding prompt engineering is an effective method in evading AI detection tools and highlights that improvements to AI detection software are needed. Using the Turnitin AI detect tool, faculty reported 54.5% of the experimental submissions to the academic misconduct process, suggesting the need for increased awareness and training into these tools. Genuine submissions received a mean score of 54.4, whereas AI-generated content scored 52.3, indicating the comparable performance of GPT-4 in real-life situations. Recommendations include adjusting assessment strategies to make them more resistant to the use of AI tools, using AI-inclusive assessment where possible, and providing comprehensive training programs for faculty and students. This research contributes to understanding the relationship between AI-generated content and academic assessment, urging further investigation to preserve academic integrity.
Motivation & Objective
- Assess robustness of university assessments against GPT-4 generated content.
- Evaluate faculty detection capability with Turnitin AI detection support.
- Quantify detection rates and impact on academic integrity indicators.
Proposed method
- Generate 22 GPT-4 submissions and embed them in an assessment.
- Have 15 faculty members mark the submissions.
- Use Turnitin AI detect tool to identify AI-generated content.
- Compare detected content to actual AI content to assess evasion.
- Analyze submission outcomes: misconduct referrals and grades.
Experimental results
Research questions
- RQ1How effective is Turnitin's AI detection at identifying GPT-4 content in real submissions?
- RQ2To what extent can adversarial prompt-engineering evade AI-detection tools in academic assessments?
- RQ3What is the impact of AI-generated content on actual assessment grades and misconduct referrals?
Key findings
- AI detection tool identified 91% of experimental AI-generated submissions as containing some AI content.
- Total detected AI content was 54.8%, indicating substantial evasion of detection despite tool flags.
- Faculty referred 54.5% of experimental submissions to misconduct processes.
- Genuine submissions averaged 54.4, while AI-generated content averaged 52.3, showing similar performance.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.