[논문 리뷰] Prescribing Deep Attentive Score Prediction Attracts Improved Student Engagement
이 연구는 스코어 예측을 위한 딥 어텐셔널 신경망 모델의 영향을 인텔리전트 튜터링 시스템(Santa)에서 평가하며, 더 높은 예측 정확도가 학생의 참여도를 크게 향상시킨다는 것을 입증한다. 약 78만 명의 사용자를 대상으로 한 A/B 테스트에서, 모델은 평균 절대 오차를 협업 필터링의 78.9에서 49.8로 감소시켰고, 이로 인해 구매율은 15.19% 상승했으며, 평균 문제 해결 수는 22.73대 20.03으로 증가했고, 총 수익은 13% 증가했다. 이는 정확한 스코어 예측이 동기부여와 능동적 참여를 모두 향상시킨다는 것을 입증한다.
Intelligent Tutoring Systems (ITSs) have been developed to provide students with personalized learning experiences by adaptively generating learning paths optimized for each individual. Within the vast scope of ITS, score prediction stands out as an area of study that enables students to construct individually realistic goals based on their current position. Via the expected score provided by the ITS, a student can instantaneously compare one's expected score to one's actual score, which directly corresponds to the reliability that the ITS can instill. In other words, refining the precision of predicted scores strictly correlates to the level of confidence that a student may have with an ITS, which will evidently ensue improved student engagement. However, previous studies have solely concentrated on improving the performance of a prediction model, largely lacking focus on the benefits generated by its practical application. In this paper, we demonstrate that the accuracy of the score prediction model deployed in a real-world setting significantly impacts user engagement by providing empirical evidence. To that end, we apply a state-of-the-art deep attentive neural network-based score prediction model to Santa, a multi-platform English ITS with approximately 780K users in South Korea that exclusively focuses on the TOEIC (Test of English for International Communications) standardized examinations. We run a controlled A/B test on the ITS with two models, respectively based on collaborative filtering and deep attentive neural networks, to verify whether the more accurate model engenders any student engagement. The results conclude that the attentive model not only induces high student morale (e.g. higher diagnostic test completion ratio, number of questions answered, etc.) but also encourages active engagement (e.g. higher purchase rate, improved total profit, etc.) on Santa.
연구 동기 및 목표
- 인텔리전트 튜터링 시스템(ITS)에서 스코어 예측 정확도 향상이 학생의 참여도에 미치는 영향을 조사하는 것.
- 최신 기술 기반의 딥 어텐셔널 신경망 모델과 전통적인 협업 필터링 간의 사용자 행동에 대한 실제 영향을 평가하는 것.
- 완료율, 구매, 수익과 같은 측정 가능한 참여 지표와의 인과관계를 정량화하는 것.
- 높은 예측 신뢰도가 교육 AI 시스템에서 사용자 신뢰와 동기부여를 증가시킨다는 경험적 증거를 제공하는 것.
제안 방법
- 한국에서 약 78만 명의 사용자를 보유한 다중 플랫폼 TOEIC 시험 대비 인텔리전트 튜터링 시스템(Santa)에서 제어된 A/B 테스트를 실시하였다.
- 두 가지 스코어 예측 모델을 배포: 협업 필터링 기반 모델(MAE = 78.9)과 딥 어텐셔널 신경망 기반 모델(MAE = 49.8).
- 딥 어텐셔널 모델은 사전 훈련/피넷유징 전략을 사용했다: 먼저 정답 여부와 응답 시간을 기반으로 사전 훈련한 후, 트랜스포머 기반 아키텍처를 사용해 시험 스코어 예측을 위한 피넷유징을 수행했다.
- 모델은 주로 학생의 학습 시퀀스에서 장거리 의존성을 포착하기 위해 어텐션 메커니즘을 활용하여 지식 상태의 표현을 향상시켰다.
- 다양한 참여 지표를 기반으로 사용자 행동을 모니터링: 진단 테스트 완료율, 멤버십 등록, 문제 해결 수, 구매율, ARPU, 총 수익.
- 모든 재무 지표는 공정한 비교를 위해 모델 파라미터 비율 기반으로 정규화되었다.
실험 결과
연구 질문
- RQ1더 정확한 스코어 예측 모델이 ITS에서 학생의 동기부여를 높이는가?
- RQ2예측 정확도 향상이 활동적 참여, 예를 들어 구매 결정에 얼마나 큰 영향을 미치는가?
- RQ3딥 어텐셔널 신경망 모델이 실제 참여 결과에서 협업 필터링을 능가하는가?
- RQ4예측 신뢰도와 사용자 유지를 비롯한 수익 창출 간에 측정 가능한 인과관계가 존재하는가?
주요 결과
- 딥 어텐셔널 모델은 평균 절대 오차(MAE)가 49.8로, 협업 필터링 모델의 MAE 78.9보다 유의미하게 낮았다.
- 딥 어텐셔널 모델의 진단 테스트 완료율은 65.90%였고, 협업 필터링 모델은 64.93%였으며, 이는 더 높은 동기부여를 시사한다.
- 딥 어텐셔널 모델의 등록(멤버십) 비율은 44.55%였고, 협업 필터링 모델은 43.13%였다.
- 딥 어テン셔널 모델 사용자는 진단 테스트 후 평균 22.73道 문제를 해결했고, 협업 필터링 모델 사용자는 20.03道였다.
- 딥 어텐셔널 모델은 구매율이 15.19% 상승했으며, 이는 2.73% 대비 2.37%로 증가했다.
- 총 수익은 딥 어텐셔널 모델 기준 162,933.88달러였고, 협업 필터링 모델 기준 142,949.55달러였으며, 정규화 후 13% 증가했다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.