Skip to main content
QUICK REVIEW

[논문 리뷰] FrontierScience: Evaluating AI's Ability to Perform Expert-Level Scientific Tasks

Miles Wang, Robi Lin|arXiv (Cornell University)|2026. 01. 29.
Machine Learning in Materials Science인용 수 0
한 줄 요약

FrontierScience는 수백 개의 전문가 수준 물리학, 화학, 생물학 문제를 통해 AI 추론을 평가하는 두 트랙 벤치마크(Olympiad 및 Research)를 도입합니다; GPT-5.2는 Olympiad에서 선두(77%), Research에서는 뒤처집니다(25%).

ABSTRACT

We introduce FrontierScience, a benchmark evaluating expert-level scientific reasoning in frontier language models. Recent model progress has nearly saturated existing science benchmarks, which often rely on multiple-choice knowledge questions or already published information. FrontierScience addresses this gap through two complementary tracks: (1) Olympiad, consisting of international olympiad problems at the level of IPhO, IChO, and IBO, and (2) Research, consisting of PhD-level, open-ended problems representative of sub-tasks in scientific research. FrontierScience contains several hundred questions (including 160 in the open-sourced gold set) covering subfields across physics, chemistry, and biology, from quantum electrodynamics to synthetic organic chemistry. All Olympiad problems are originally produced by international Olympiad medalists and national team coaches to ensure standards of difficulty, originality, and factuality. All Research problems are research sub-tasks written and verified by PhD scientists (doctoral candidates, postdoctoral researchers, or professors). For Research, we introduce a granular rubric-based evaluation framework to assess model capabilities throughout the process of solving a research task, rather than judging only a standalone final answer.

연구 동기 및 목표

  • 제한된(Olympiad) 과제와 개방형(Open-ended) 연구 과제 전반에 걸쳐 전문가 수준의 과학적 추론 능력에서 AI의 역량을 평가한다.
  • 난이도와 독창성을 보장하기 위해 도메인 전문가가 확인한 신규의 전문가가 저술한 문제를 제공한다.
  • 모델의 강점과 약점을 진단하기 위한 개방형 연구 과제용 루브릭 기반 평가 프레임워크를 도입한다.

제안 방법

  • 두 트랙 데이터셋: FrontierScience-Olympiad는 짧은 답변과 문제 해결 문제로 구성; FrontierScience-Research는 PhD 수준의 개방형 하위 문제로 구성.
  • 물리학, 화학, 생물학의 도메인 전문가가 작성하고 검증한 문제들; 각 Research 문제에는 10점 루브릭과 설명 솔루션 경로가 포함.
  • 개방형 연구 과제에 대한 루브릭 기반 채점으로 중간 추론 및 최종 답안을 평가하며, 채점은 모델 주관자(GPT-5)를 사용.
  • 고도 추론 노력을 요하는 다중 프런티어 모델을 사용한 평가를 진행하며 Olympiad에 20회의 시험, Research에 30회의 시험을 수행; 모델 판단은 GPT-5 기반 심판이 수행.
  • 메타 리뷰 및 더 큰 말뭄에서 필터링 후 100개의 Olympiad 문제와 60개의 Research 문제가 포함된 오픈 소스 골드 세트.
Figure 1: Sample FrontierScience-Olympiad problems. Each task in FrontierScience is written and verified by a domain expert in physics, chemistry, or biology. For the Olympiad set, all experts achieved a medal in an international olympiad competition.
Figure 1: Sample FrontierScience-Olympiad problems. Each task in FrontierScience is written and verified by a domain expert in physics, chemistry, or biology. For the Olympiad set, all experts achieved a medal in an international olympiad competition.

실험 결과

연구 질문

  • RQ1제한된 형태의 해가 닫히거나 숫자/표현 가능한 답이 있는 Olympiad 스타일의 물리학, 화학, 생물학 문제를 프런티어 AI 모델이 얼마나 잘 해결하나요?
  • RQ2추론, 정당화, 루브릭 기반 평가가 필요한 개방형 PhD 수준 연구 하위 문제를 프런티어 AI 모델이 얼마나 잘 다루나요?
  • RQ3제한된 과제와 개방형 과학 과제 전반에서 현재 프런티어 모델의 강점과 실패 모드는 무엇인가요?
  • RQ4각 트랙 내에서 물리학, 화학, 생물학에 따른 모델 성능 차이는 어떻게 나타나나요?

주요 결과

  • GPT-5.2는 FrontierScience에서 테스트된 모델 중 전체 성능이 가장 높으며 Olympiad에서 77%, Research에서 25%를 달성했습니다.
  • Gemini 3 Pro는 Olympiad 문제에서 GPT-5.2와 비슷한 수준(76%)이고, Research 세트에서 GPT-5가 GPT-5.2와 동률(25%)입니다.
  • 문제 전반에서 모델은 Olympiad 문제의 경우 화학이 가장 높은 성능을 보이고 그다음 물리학, 생물학 순; Research에서는 화학이 주도하고 생물학과 물리학이 그 뒤를 잇습니다.
  • 테스트 시간 토큰을 늘리면 GPT-5.2의 성능이 향상됩니다( Olympiad: 67.5%에서 77.1%로; Research: 18%에서 25%로 ).
  • 평가 파이프라인은 Research 과제에 대해 루브릭 기반 채점 아키텍처를 사용하고 Olympiad 과제에는 수치/표현 매칭을 사용하며, GPT-5 기반 심판이 루브릭 충족도를 평가합니다.
Figure 2: Sample FrontierScience-Research problems. For the Research set, all experts hold a relevant PhD degree. The corresponding rubrics to these sample tasks can be found in Appendix A .
Figure 2: Sample FrontierScience-Research problems. For the Research set, all experts hold a relevant PhD degree. The corresponding rubrics to these sample tasks can be found in Appendix A .

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.