Skip to main content
QUICK REVIEW

[论文解读] FrontierScience: Evaluating AI's Ability to Perform Expert-Level Scientific Tasks

Miles Wang, Robi Lin|arXiv (Cornell University)|Jan 29, 2026
Machine Learning in Materials Science被引用 0
一句话总结

FrontierScience 引入两轨基准测试(奥林匹亚与研究)含数百道物理、化学与生物学专家级题目以评估 AI 推理;GPT-5.2 在奥林匹亚(77%)领先,在研究(25%)落后。

ABSTRACT

We introduce FrontierScience, a benchmark evaluating expert-level scientific reasoning in frontier language models. Recent model progress has nearly saturated existing science benchmarks, which often rely on multiple-choice knowledge questions or already published information. FrontierScience addresses this gap through two complementary tracks: (1) Olympiad, consisting of international olympiad problems at the level of IPhO, IChO, and IBO, and (2) Research, consisting of PhD-level, open-ended problems representative of sub-tasks in scientific research. FrontierScience contains several hundred questions (including 160 in the open-sourced gold set) covering subfields across physics, chemistry, and biology, from quantum electrodynamics to synthetic organic chemistry. All Olympiad problems are originally produced by international Olympiad medalists and national team coaches to ensure standards of difficulty, originality, and factuality. All Research problems are research sub-tasks written and verified by PhD scientists (doctoral candidates, postdoctoral researchers, or professors). For Research, we introduce a granular rubric-based evaluation framework to assess model capabilities throughout the process of solving a research task, rather than judging only a standalone final answer.

研究动机与目标

  • 评估 AI 在受限(奥林匹亚)与开放式(研究)任务中的专家级科学推理能力。
  • 提供新颖、专家撰写并经领域专家验证的题目,以确保难度与原创性。
  • 引入基于评分的开放式研究任务评估框架,以诊断模型的优点与弱点。

提出的方法

  • 两轨数据集:FrontierScience-Olympiad,包含简答题与解题题;FrontierScience-Research,包含博士级别的开放式子问题。
  • 题目由物理、化学与生物领域专家撰写并验证;每道研究题包含一个10分制的评分标准及解释性解题路径。
  • 对研究任务采用基于评分的评定以评估中间推理与最终答案,使用基于 GPT-5 的裁判进行评分。
  • 评估在高推理强度下对多种前沿模型进行,奥林匹亚阶段进行20次试验,研究阶段30次试验;模型裁判由 GPT-5 为基础的裁判执行。
  • 开源金标准集:在元评审与从更大语料库筛选后,包含100道奥林匹亚题和60道研究题。
Figure 1: Sample FrontierScience-Olympiad problems. Each task in FrontierScience is written and verified by a domain expert in physics, chemistry, or biology. For the Olympiad set, all experts achieved a medal in an international olympiad competition.
Figure 1: Sample FrontierScience-Olympiad problems. Each task in FrontierScience is written and verified by a domain expert in physics, chemistry, or biology. For the Olympiad set, all experts achieved a medal in an international olympiad competition.

实验结果

研究问题

  • RQ1前沿 AI 模型在具有闭式解或可表达数值/公式答案的奥林匹亚风格物理、化学与生物学题目上能否较好解决?
  • RQ2前沿 AI 模型在需要推理、论证和基于评分的评估的博士级开放式研究子问题上能否较好解决?
  • RQ3当前前沿模型在受限 versus 开放式科学任务中的优势与失败模式是什么?
  • RQ4在各轨道中模型在不同学科(物理、化学、生物)上的表现如何变化?

主要发现

  • GPT-5.2 在 FrontierScience 测试中的总体表现最高,奥林匹亚题为 77%,研究题为 25%。
  • Gemini 3 Pro 在奥林匹亚题上可与 GPT-5.2 相提并论(76%),GPT-5 在研究集(25%)与 GPT-5.2 打成平手。
  • 在所有题目中,模型在奥林匹亚题对化学的表现最强,其次是物理,生物学最弱;在研究题中,化学领先,其次是生物学与物理学。
  • 更多测试时标记的代币提高 GPT-5.2 的表现(奥林匹亚:67.5% 提升至 77.1%;研究:18% 提升至 25%)。
  • 评估流程对研究任务使用基于评分的结构化评分架构,对奥林匹亚任务使用数值/表达式匹配;由 GPT-5 作为裁判评估评分完成情况。
Figure 2: Sample FrontierScience-Research problems. For the Research set, all experts hold a relevant PhD degree. The corresponding rubrics to these sample tasks can be found in Appendix A .
Figure 2: Sample FrontierScience-Research problems. For the Research set, all experts hold a relevant PhD degree. The corresponding rubrics to these sample tasks can be found in Appendix A .

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。