[논문 리뷰] Evaluating GPT-3.5 and GPT-4 Models on Brazilian University Admission Exams
이 논문은 포르투갈어로 된 브라질 ENEM 시험에서 GPT-3.5와 GPT-4를 평가하며, Chain-of-Thought 프롬프트를 사용한 GPT-4가 ENEM 2022에서 최대 87%의 정확도를 달성하고, GPT-3.5보다 약 11포인트 앞서며, CoT가 수학 및 과학 문제에서 상당한 향상을 가능하게 한다.
The present study aims to explore the capabilities of Language Models (LMs) in tackling high-stakes multiple-choice tests, represented here by the Exame Nacional do Ensino Médio (ENEM), a multidisciplinary entrance examination widely adopted by Brazilian universities. This exam poses challenging tasks for LMs, since its questions may span into multiple fields of knowledge, requiring understanding of information from diverse domains. For instance, a question may require comprehension of both statistics and biology to be solved. This work analyzed responses generated by GPT-3.5 and GPT-4 models for questions presented in the 2009-2017 exams, as well as for questions of the 2022 exam, which were made public after the training of the models was completed. Furthermore, different prompt strategies were tested, including the use of Chain-of-Thought (CoT) prompts to generate explanations for answers. On the 2022 edition, the best-performing model, GPT-4 with CoT, achieved an accuracy of 87%, largely surpassing GPT-3.5 by 11 points. The code and data used on experiments are available at https://github.com/piresramon/gpt-4-enem.
연구 동기 및 목표
- 최첨단 언어 모델이 브라질 포르투갈어의 고위험 다영역 시험인 ENEM에서 어떻게 수행하는지 평가합니다.
제안 방법
- 제로샷, 파샷 및 Chain-of-Thought 프롬프트를 통한 GPT-3.5 및 GPT-4 변형을 평가합니다.

실험 결과
연구 질문
- RQ1GPT-3.5와 GPT-4가 ENEM 2009-2017 및 ENEM 2022 데이터셋에서 얼마나 잘 작동합니까?
- RQ2Chain-of-Thought 프롬프팅이 다영역의 포르투갈어 시험 문제에서 정확도를 향상시키나요?
- RQ3프롬프트가 모델의 2022 ENEM 콘텐츠 암기 가능성을 드러내거나 완화시킬 수 있나요?
주요 결과
- CoT를 사용하는 GPT-4는 ENEM 2022에서 87%의 정확도를 달성하여 GPT-3.5보다 약 11포인트 앞섭니다.

더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.