Skip to main content
QUICK REVIEW

[Paper Review] SciKnowEval: Evaluating Multi-level Scientific Knowledge of Large Language Models

Kehua Feng, Shen, Xinyi|arXiv (Cornell University)|Jun 13, 2024
Topic Modeling4 citations
TL;DR

SciKnowEval introduces a comprehensive benchmark to evaluate large language models (LLMs) across five progressive levels of scientific knowledge—studying extensively, inquiring earnestly, thinking profoundly, discerning clearly, and practicing assiduously—using a 50K-sample dataset in biology and chemistry. Despite state-of-the-art performance, even top models like GPT-4o show significant gaps in scientific reasoning, safety awareness, and experimental protocol generation, highlighting critical limitations in real-world scientific application.

ABSTRACT

Large language models (LLMs) are playing an increasingly important role in scientific research, yet there remains a lack of comprehensive benchmarks to evaluate the breadth and depth of scientific knowledge embedded in these models. To address this gap, we introduce SciKnowEval, a large-scale dataset designed to systematically assess LLMs across five progressive levels of scientific understanding: memory, comprehension, reasoning, discernment, and application. SciKnowEval comprises 28K multi-level questions and solutions spanning biology, chemistry, physics, and materials science. Using this benchmark, we evaluate 20 leading open-source and proprietary LLMs. The results show that while proprietary models often achieve state-of-the-art performance, substantial challenges remain -- particularly in scientific reasoning and real-world application. We envision SciKnowEval as a standard benchmark for evaluating scientific capabilities in LLMs and as a catalyst for advancing more capable and reliable scientific language models.

Motivation & Objective

  • To address the lack of comprehensive, multi-level benchmarks for evaluating scientific knowledge in LLMs.
  • To overcome limitations in existing benchmarks that focus only on high school-level science or lack ethical and safety evaluation.
  • To develop a systematic framework inspired by Confucian philosophy that mirrors the human scientific learning process.
  • To assess LLMs across five progressive levels: knowledge acquisition, inquiry, reasoning, ethical discernment, and practical application.
  • To promote the development of safer, more capable scientific LLMs through a large-scale, diverse, and ethically informed benchmark.

Proposed method

  • Design a five-level scientific knowledge evaluation framework inspired by the 'Doctrine of the Mean', modeling the progression from knowledge acquisition to practical application.
  • Construct a large-scale, multi-source dataset of 50,048 scientific problems and solutions from textbooks, literature, databases, and self-generated content in biology and chemistry.
  • Categorize tasks into five levels: Studying Extensively (knowledge recall), Enquiring Earnestly (inquiry), Thinking Profoundly (reasoning), Discerning Clearly (ethics/safety), and Practicing Assiduously (experimental design).
  • Use zero-shot and few-shot prompting to evaluate 20 leading open-source and proprietary LLMs across all five levels.
  • Incorporate human-annotated scoring (e.g., GPT-4o as a judge) to assess model outputs on accuracy, completeness, and safety in complex scientific tasks.
  • Ensure ethical evaluation by including safety-critical tasks such as hazardous chemical synthesis and regulatory compliance.

Experimental results

Research questions

  • RQ1How well do LLMs perform across a multi-level hierarchy of scientific knowledge, from basic recall to complex experimental design?
  • RQ2To what extent do proprietary and open-source LLMs differ in reasoning, safety awareness, and practical application of scientific knowledge?
  • RQ3Can existing benchmarks adequately capture the depth of scientific understanding required for real-world research and experimentation?
  • RQ4How do LLMs handle ethical and safety considerations in high-risk scientific tasks, such as synthesizing banned chemicals?
  • RQ5What are the critical failure points in LLM-generated scientific protocols, especially in reagent selection and procedural detail?

Key findings

  • Even the most advanced LLM, GPT-4o, failed to achieve a score of 3/5 on average in reagent and procedure generation tasks at the highest (L5) level, indicating significant gaps in practical scientific expertise.
  • Proprietary LLMs, despite state-of-the-art performance, show notable deficiencies in scientific computation and application, particularly in handling safety-critical tasks.
  • No model fully mastered the L5 level (Practicing Assiduously), with GPT-4o scoring only 2/5 on a lipidomics SPE protocol design task, missing key reagents and dosages.
  • The benchmark reveals that current LLMs often lack specificity and completeness in experimental design, especially in naming correct solvents, ratios, and internal standards.
  • Ethical and safety considerations are consistently under-evaluated in existing benchmarks, but SciKnowEval explicitly includes them, exposing a critical blind spot in current LLM evaluation.
  • The dataset and code are publicly available at https://github.com/hicai-zju/sciknoweval, enabling reproducible and standardized evaluation of scientific LLMs.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.