Skip to main content
QUICK REVIEW

[Paper Review] Agentic AI for Scientific Discovery: A Survey of Progress, Challenges, and Future Directions

Mourad Gridach, Jay Nanavati|ArXiv.org|Mar 12, 2025
Scientific Computing and Data Management10 citations
TL;DR

This survey reviews agentic AI systems for scientific discovery, outlining architectures (autonomous and collaborative), literature-review hurdles, datasets, metrics, challenges, and future directions. It covers chemistry, biology, materials science, and beyond.

ABSTRACT

The integration of Agentic AI into scientific discovery marks a new frontier in research automation. These AI systems, capable of reasoning, planning, and autonomous decision-making, are transforming how scientists perform literature review, generate hypotheses, conduct experiments, and analyze results. This survey provides a comprehensive overview of Agentic AI for scientific discovery, categorizing existing systems and tools, and highlighting recent progress across fields such as chemistry, biology, and materials science. We discuss key evaluation metrics, implementation frameworks, and commonly used datasets to offer a detailed understanding of the current state of the field. Finally, we address critical challenges, such as literature review automation, system reliability, and ethical concerns, while outlining future research directions that emphasize human-AI collaboration and enhanced system calibration.

Motivation & Objective

  • Define agentic AI and its role in accelerating scientific discovery.
  • Categorize autonomous and collaborative agentic AI systems and their domain applications.
  • Identify datasets, tools, and evaluation metrics used in the field.
  • Highlight challenges in literature review automation, reliability, and ethics.
  • Propose future directions emphasizing human-AI collaboration and system calibration.

Proposed method

  • Review and synthesis of existing agentic AI systems and frameworks (e.g., Coscientist, ChemCrow, ProtAgents, LLaMP, Organa).
  • Taxonomy development for fully autonomous vs. human-AI collaborative systems.
  • Assessment of literature-review frameworks (SciLitLLM, LitSearch, ResearchArena, CiteME) and their limitations.
  • Compilation of implementation tools (AutoGen, MetaGPT, Letta, CAMEL, LangChain, AutoGPT) and datasets (LAB-Bench, MoleculeNet, ZINC, MPcules, AlphaFold, PubChem, ChEMBL).
  • Discussion of evaluation metrics including NeurIPS-style paper evaluation, success rates, and usability metrics.
  • Synthesis of challenges (trustworthiness, ethics, potential risks) and future directions (calibration, governance).

Experimental results

Research questions

  • RQ1What are the current architectures and frameworks for agentic AI in scientific discovery?
  • RQ2What datasets, tools, and metrics support evaluation and benchmarking of agentic AI systems?
  • RQ3What are the main challenges (literature review, reliability, ethics) limiting system performance and adoption?
  • RQ4How can human-AI collaboration and calibration improve the trustworthiness and impact of agentic AI in science?

Key findings

  • Agentic AI systems can automate stages from ideation to paper writing, showing progress in chemistry, biology, and materials science.
  • Literature review remains a major bottleneck across frameworks, with several approaches failing to reliably ground ideas in existing knowledge.
  • Multiple autonomous and collaborative frameworks exist, each with domain strengths and limitations in interpretability and generalizability.
  • A range of tools (AutoGen, LangChain, Letta, CAMEL) and datasets (LAB-Bench, MoleculeNet, ZINC, AlphaFold) are used to develop and evaluate these agents.
  • Evaluation frameworks are evolving, with NeurIPS-style paper evaluation and usability metrics increasingly incorporated to assess scientific outputs and system usefulness.
  • Future directions emphasize human-in-the-loop approaches, better calibration techniques, and robust governance to ensure reliable and ethical autonomous scientific discovery.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.