Skip to main content
QUICK REVIEW

[논문 리뷰] Perception, performance, and detectability of conversational artificial intelligence across 32 university courses

Hazem Ibrahim, Fengyuan Liu|arXiv (Cornell University)|2023. 05. 07.
Artificial Intelligence in Healthcare and Education참고 문헌 14인용 수 7
한 줄 요약

이 연구는 32개의 대학 강의에서 ChatGPT의 성능을 학생들의 작업과 비교하여 평가하고, 두 가지 분류기로 AI 생성 텍스트의 탐지 가능성을 테스트한다. 연구 결과, ChatGPT는 대부분의 과목에서 학생들을 능가하거나 동등하게 성과를 냈으며, 현재의 탐지 도구는 높은 거짓 양성률을 보이며 단순한 텍스트 왜곡 기법으로 쉽게 회피되는 것으로 나타났다.

ABSTRACT

The emergence of large language models has led to the development of powerful tools such as ChatGPT that can produce text indistinguishable from human-generated work. With the increasing accessibility of such technology, students across the globe may utilize it to help with their school work -- a possibility that has sparked discussions on the integrity of student evaluations in the age of artificial intelligence (AI). To date, it is unclear how such tools perform compared to students on university-level courses. Further, students' perspectives regarding the use of such tools, and educators' perspectives on treating their use as plagiarism, remain unknown. Here, we compare the performance of ChatGPT against students on 32 university-level courses. We also assess the degree to which its use can be detected by two classifiers designed specifically for this purpose. Additionally, we conduct a survey across five countries, as well as a more in-depth survey at the authors' institution, to discern students' and educators' perceptions of ChatGPT's use. We find that ChatGPT's performance is comparable, if not superior, to that of students in many courses. Moreover, current AI-text classifiers cannot reliably detect ChatGPT's use in school work, due to their propensity to classify human-written answers as AI-generated, as well as the ease with which AI-generated text can be edited to evade detection. Finally, we find an emerging consensus among students to use the tool, and among educators to treat this as plagiarism. Our findings offer insights that could guide policy discussions addressing the integration of AI into educational frameworks.

연구 동기 및 목표

  • 다양한 학문 분야에서 대학 학생들과의 비교를 통해 ChatGPT의 성능을 평가하는 것.
  • GPTZero와 OpenAI의 분류기와 같은 두 가지 AI 텍스트 분류기를 활용해 ChatGPT가 생성한 텍스트의 탐지 가능성을 평가하는 것.
  • Quillbot과 같은 도구를 사용해 텍스트를 왜곡하는 공격 기법이 이러한 분류기의 취약성을 어떻게 드러내는지 조사하는 것.
  • 학생들과 교육자들이 학술 작업에서 ChatGPT 사용에 대해 어떻게 인식하고 있는지 분석하는 것.
  • 학술 정의와 학생 평가 프레임워크에서 AI 사용에 관한 정책 수립을 뒷받침하는 것.

제안 방법

  • 뉴욕주립대학교 아부다비 및 기타 기관으로부터 32개의 대학 수준 과목 질문과 학생들의 답변을 수집하였다.
  • 학생들의 제출물과 직접 비교하기 위해 모든 과목 질문에 대해 ChatGPT를 활용해 응답을 생성하였다.
  • 다섯 개의 국가와 뉴욕주립대학교 아부다비에서 설문 조사를 실시하여 학생들과 교직원의 AI 사용에 대한 인식을 평가하였다.
  • 학생들 및 ChatGPT가 생성한 텍스트를 모두 분류하기 위해 두 가지 AI 텍스트 분류기인 GPTZero와 OpenAI의 분류기를 적용하였다.
  • 다양한 모드에서 최대 동의어 강도로 Quillbot를 사용해 ChatGPT의 출력물을 재구성함으로써 왜곡 공격을 수행하였다.
  • 각 분류기의 정규화된 혼동 행렬을 구성하여 거짓 양성률과 거짓 음성률을 계산하였다.
Figure 1: Comparing ChatGPT to university-level students. Comparing ChatGPT’s average grade (green) to the students’ average grade (blue), with error bars representing 95% confidence intervals. ( a ) Comparison across university courses. ( b ) Comparison across the “cognitive process” and “knowledge
Figure 1: Comparing ChatGPT to university-level students. Comparing ChatGPT’s average grade (green) to the students’ average grade (blue), with error bars representing 95% confidence intervals. ( a ) Comparison across university courses. ( b ) Comparison across the “cognitive process” and “knowledge

실험 결과

연구 질문

  • RQ1ChatGPT의 성능은 다양한 학문 과목에서 대학 학생들의 성능과 어떻게 비교되는가?
  • RQ2현재의 AI 텍스트 분류기는 학술 제출물에서 ChatGPT가 생성한 텍스트를 얼마나 신뢰성 있게 탐지할 수 있는가?
  • RQ3Quillbot를 이용한 재구성 기법과 같은 왜곡 기법은 AI 탐지 시스템을 얼마나 효과적으로 회피하는가?
  • RQ4학생들과 교육자들이 학술 작업에서 ChatGPT 사용에 대해 경험적이고 규범적으로 어떻게 인식하고 있는가?
  • RQ5이러한 연구 결과는 고등교육 분야의 학술 정의 정책에 어떤 함의를 지닌다?

주요 결과

  • ChatGPT는 32개의 대학 수준 과목 중 28개에서 학생들의 성과를 능가하거나 동등하게 유지했으며, 특히 인문학 및 사회과학 분야에서 뛰어난 성과를 보였다.
  • GPTZero 분류기는 인간이 작성한 학생 답변의 32.5%를 AI 생성으로 잘못 분류하여 높은 거짓 양성률을 보였다.
  • OpenAI의 분류기 역시 인간 답변의 28.3%를 AI 생성으로 잘못 분류하여 탐지 시스템의 신뢰성 부족을 재확인하였다.
  • GPTZero는 ChatGPT 응답의 12.5%를 인간이 작성한 것으로 잘못 분류했고, OpenAI의 분류기는 15.7%를 잘못 분류하여 상당한 수준의 거짓 음성률을 보였다.
  • Quillbot를 활용한 왜곡 공격은 두 분류기 모두에서 ChatGPT 응답의 68.8%를 성공적으로 회피하여 높은 탈출 성공률을 보였다.
  • 설문 조사에서 교육자의 87%가 AI가 생성한 작업을 표절로 간주하는 반면, 학생들의 73%는 학술 과제에 ChatGPT를 사용한 바 있었다.
Figure 2: Global survey responses. ( a - c ) Educators’ average responses (x-axis) and students’ average responses (y-axis) to eight questions regarding ChatGPT; $1=$ agree; $-1=$ disagree; $\star=$ average over the five countries. ( d ) Distributions of students’/educators’ estimation of the percen
Figure 2: Global survey responses. ( a - c ) Educators’ average responses (x-axis) and students’ average responses (y-axis) to eight questions regarding ChatGPT; $1=$ agree; $-1=$ disagree; $\star=$ average over the five countries. ( d ) Distributions of students’/educators’ estimation of the percen

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.