Skip to main content
QUICK REVIEW

[논문 리뷰] Can AI grade your essays? A comparative analysis of large language models and teacher ratings in multidimensional essay scoring

Kathrin Seßler, Maurice Fürstenberg|arXiv (Cornell University)|2024. 11. 25.
Online Learning and Analytics인용 수 4
한 줄 요약

이 연구는 독일 교육 분야에서 다차원 에세이 점수 평가를 위한 오픈소스 및 비공개소스 대규모 언어 모델(Large Language Models, LLMs)을 평가하며, 10개 기준에 따라 37名의 교사 평가와 비교한다. 신규 모델인 o1은 높은 신뢰성(ICC = .80)과 강한 상관관계(Spearman의 r = .74)를 보이며 인간 평가자와 유사한 성능을 보이며, 언어 관련 기준에서는 특히 뛰어난 성능을 보였지만 전체 점수를 높게 부여하는 경향이 있어 내용 품질 평가의 개선이 필요함을 시사한다.

ABSTRACT

The manual assessment and grading of student writing is a time-consuming yet critical task for teachers. Recent developments in generative AI, such as large language models, offer potential solutions to facilitate essay-scoring tasks for teachers. In our study, we evaluate the performance and reliability of both open-source and closed-source LLMs in assessing German student essays, comparing their evaluations to those of 37 teachers across 10 pre-defined criteria (i.e., plot logic, expression). A corpus of 20 real-world essays from Year 7 and 8 students was analyzed using five LLMs: GPT-3.5, GPT-4, o1, LLaMA 3-70B, and Mixtral 8x7B, aiming to provide in-depth insights into LLMs' scoring capabilities. Closed-source GPT models outperform open-source models in both internal consistency and alignment with human ratings, particularly excelling in language-related criteria. The novel o1 model outperforms all other LLMs, achieving Spearman's $r = .74$ with human assessments in the overall score, and an internal consistency of $ICC=.80$. These findings indicate that LLM-based assessment can be a useful tool to reduce teacher workload by supporting the evaluation of essays, especially with regard to language-related criteria. However, due to their tendency for higher scores, the models require further refinement to better capture aspects of content quality.

연구 동기 및 목표

  • 다양한 차원에서 독일 학생 에세이 점수 평가에 있어 대규모 언어 모델(Large Language Models, LLMs)의 신뢰성과 타당성을 평가하기 위해.
  • 사전 정의된 10개 기준에서 오픈소스 및 비공개소스 LLM의 성능을 인간 교사 평가와 비교하기 위해.
  • 교육적 에세이 점수 평가에서 LLM의 내재 일관성과 인간 평가와의 일치도를 평가하기 위해.
  • 언어 품질 평가 대비 내용 품질 평가에서 LLM의 강점과 한계를 파악하기 위해, 특히 교사 지원에 있어.
  • 교사의 부담을 줄이면서도 평가 품질을 유지할 수 있도록 AI 지원 에세이 점수 평가 도구의 개발을 안내하기 위해.

제안 방법

  • 7학년과 8학년 학생들의 실제 에세이 20편으로 구성된 코퍼스를 대상으로, GPT-3.5, GPT-4, o1, LLaMA 3-70B, Mixtral 8x7B 총 5종의 LLM을 사용하여 평가하였다.
  • 에세이는 문학적 논리, 표현 방식, 언어 정확도 등 10개의 사전 정의된 기준에 따라 37명의 교사가 평가하였다.
  • 모델의 행동을 고립하고 일관성을 확보하기 위해 프롬프트 엔지니어링 없이 LLM 평가를 수행하였다.
  • 정렬 상관관계(Spearman’s rank correlation, r)와 집단간 상관계수(Intraclass Correlation, ICC)와 같은 통계적 지표를 사용하여 일치도와 내재 일관성을 평가하였다.
  • 비공개소스 모델(GPT-4, o1 등)과 오픈소스 모델(LLaMA 3, Mixtral 등) 간의 성능 차이를 중심으로 비교 분석을 수행하였다.
  • 자동화된 점수 평가의 신뢰성과 강건성을 평가하기 위해 다수의 LLM 실행 결과에 대한 분산 분석을 수행하였다.
Figure 1. Design and workflow of our study: student essays are evaluated based on predefined criteria by both Large Language Models and human teachers, followed by a comprehensive analysis of the resulting ratings.
Figure 1. Design and workflow of our study: student essays are evaluated based on predefined criteria by both Large Language Models and human teachers, followed by a comprehensive analysis of the resulting ratings.

실험 결과

연구 질문

  • RQ1비공개소스 및 오픈소스 LLM은 독일 에세이 점수 평가에서 10개의 다차원 기준에 걸쳐 인간 교사 평가와 얼마나 유사한가?
  • RQ2LLM이 생성한 점수의 내재 일관성은 어떠한가? 그리고 다양한 모델 간에 어떻게 다름이 있는가?
  • RQ3언어 관련 기준과 내용 관련 기준에서 인간 교사 평가를 가장 잘 재현하는 LLM은 무엇인가?
  • RQ4LLM은 전체 점수를 높게 부여하는 경향이 얼마나 강한가? 이는 교육 현장에서의 신뢰성에 어떤 영향을 미치는가?
  • RQ5LLM은 어떻게 개선되어야 내용 품질 평가를 더 잘 수행할 수 있을까? 동시에 인간 평가 기준과의 일치도를 유지할 수 있는가?

주요 결과

  • o1 모델은 인간 평가자와 가장 높은 상관관계를 보였으며(Spearman의 r = .74), 내재 일관성도 가장 높았다(ICC = .80)로 다른 모든 모델보다 뛰어났다.
  • 비공개소스 모델, 특히 GPT-4와 o1은 인간 평가자와의 일치도와 내재 일관성 측면에서 오픈소스 모델보다 뚜렷이 뛰어났다.
  • LLaMA 3-70B와 Mixtral 8x7B와 같은 오픈소스 모델은 분산은 낮았지만 인간 평가자와의 상관관계가 약해 다차원 평가에 있어 낮은 신뢰성을 보였다.
  • LLM은 언어 관련 기준(예: 표현, 문법)에서 내용 관련 기준(예: 서사 논리, 논거 구성)보다 더 뛰어난 성능을 보였다.
  • o1 모델의 성능은 최신의 추론 최적화된 LLM이 인간 평가를 밀도 있게 모방할 수 있음을 시사하지만, 여전히 전체 점수를 높게 부여하는 경향이 있어 내용 품질 평가의 정교화가 필요하다.
  • 본 연구는 LLM의 내용 품질 평가 능력 향상과 점수 편향 감소를 통해 인간 기준과의 일치도를 확보하기 위한 향후 개선의 필요성을 강조한다.
Figure 3. Correlation between o1 and human ratings across all evaluation criteria. The red line represents the linear least-squares regression between the two sets of ratings.
Figure 3. Correlation between o1 and human ratings across all evaluation criteria. The red line represents the linear least-squares regression between the two sets of ratings.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.