Skip to main content
QUICK REVIEW

[논문 리뷰] Exploring Durham University Physics exams with Large Language Models

Will Yeadon, D. P. Halliday|arXiv (Cornell University)|2023. 06. 27.
Artificial Intelligence in Healthcare and Education인용 수 8
한 줄 요약

GPT-4와 GPT-3.5가 더럼 대학교 물리학 기출 42개 시험(593문제, 2504점)을 평가하여 AI 능력과 시험 무결성을 확인했습니다; GPT-4의 평균은 49.4%, GPT-3.5의 평균은 38.6%였으며, COVID 이후 소폭 하락이 나타났습니다.

ABSTRACT

The emergence of advanced Natural Language Processing (NLP) models like ChatGPT has raised concerns among universities regarding AI-driven exam completion. This paper provides a comprehensive evaluation of the proficiency of GPT-4 and GPT-3.5 in answering a set of 42 exam papers derived from 10 distinct physics courses, administered at Durham University over the span of 2018 to 2022, totalling 593 questions and 2504 available marks. These exams, spanning both undergraduate and postgraduate levels, include traditional pre-COVID and adaptive COVID-era formats. Questions from the years 2018-2020 were designed for pre-COVID in person adjudicated examinations whereas the 2021-2022 exams were set for varying COVID-adapted conditions including open-book conditions. To ensure a fair evaluation of AI performances, the exams completed by AI were assessed by the original exam markers. However, due to staffing constraints, only the aforementioned 593 out of the total 1280 questions were marked. GPT-4 and GPT-3.5 scored an average of 49.4\% and 38.6\%, respectively, suggesting only the weaker students would potential improve their marks if using AI. For exams from the pre-COVID era, the average scores for GPT-4 and GPT-3.5 were 50.8\% and 41.6\%, respectively. However, post-COVID, these dropped to 47.5\% and 33.6\%. Thus contrary to expectations, the change to less fact-based questions in the COVID era did not significantly impact AI performance for the state-of-the-art models such as GPT-4. These findings suggest that while current AI models struggle with university-level Physics questions, an improving trend is observable. The code used for automated AI completion is made publicly available for further research.

연구 동기 및 목표

  • 대학 물리학에서 AI 보조 시험 완성의 위험을 동기 부여하고 정량화합니다.
  • 2018–2022년의 실제 더럼 물리학 시험에서 최첨단 LLM(GPT-4 및 GPT-3.5)의 성능을 평가합니다.
  • 복제 가능하고 투명한 방법론과 재현 가능한 오픈 소스 도구를 제공합니다.

제안 방법

  • 강의 스타일의 LaTeX 원본 파일에서 개별 문제를 정규표현식으로 자동 추출합니다.
  • 입력의 컴파일 가능성을 보장하기 위해 GPT-3.5를 사용하여 정리 및 LaTeX 오류 수정합니다.
  • 시스템 프롬트를 통해 물리학 교수 직위를 가정하고 LaTeX 형식의 답안을 작성하도록 OpenAI API에 문제를 전송합니다.
  • AI 출력물을 시험별 PDF로 편집하고 원래 코스 채점자에 의해 채점되도록 합니다.
  • 최대 세 차례 재시도를 포함한 반복적 LaTeX 컴파일 검사를 수행하고, 컴파일 실패 및 문제별 접근 이슈를 기록합니다.
  • 스크립트의 신뢰성을 보장하기 위한 추출 문제와 답안의 수동 검증을 수행하고 재현성을 위해 GitHub에 코드를 공유합니다.

실험 결과

연구 질문

  • RQ1GPT-4와 GPT-3.5가 여러 과목 및 수준에 걸쳐 더럼 대학교 물리학 시험에서 비중이 큰 점수를 달성하는가?
  • RQ2COVID 이전(대면)과 COVID 이후(개방형 책/원격 적응) 시험 형식 간 AI 성능 차이가 있는가?
  • RQ3AI 성능은 시험 수준(Levels 1–4)이나 과목 유형에 따라 달라지는가?
  • RQ4그래픽의 존재, 설명 요청, 수학적 언어 사용 등과 같은 요인이 AI 점수에 어떤 상관 관계를 보이는가?

주요 결과

  • GPT-4는 593문제에 대해 평균 49.4%, GPT-3.5는 38.6%를 기록했습니다.
  • COVID 이전 평균은 GPT-4 50.8%, GPT-3.5 41.6%였습니다.
  • COVID 이후 평균은 GPT-4 47.5%, GPT-3.5 33.6%였습니다.
  • GPT-4는 모든 시험 유형에서 GPT-3.5를 앞섰으며, Foundations of Physics 3A와 Theoretical Astrophysics에서 더 근소한 차이를 보였습니다.
  • 제로 스코어를 제외하면 비-제로 시도에서 AI 성능이 GPT-4 65.6%, GPT-3.5 56.7%로 상승합니다.
  • 이 연구는 재현을 위한 오픈 소스 코드를 제공하며 모델이 개선됨에 따라 AI 위험에 대한 지속적인 평가의 중요성을 강조합니다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.