Skip to main content
QUICK REVIEW

[Paper Review] From Intent to Evidence: A Categorical Approach for Structural Evaluation of Deep Research Agents

Shuoling Liu, Zhiquan Tan|arXiv (Cornell University)|Mar 26, 2026
Machine Learning in Materials Science0 citations
TL;DR

The paper formalizes Deep Research Agents (DRAs) via category theory and introduces a 296-question benchmark to stress-test DRA structure-preserving capabilities across four axes, revealing strong limitations in multi-hop structural synthesis.

ABSTRACT

Although deep research agents (DRAs) have emerged as a promising paradigm for complex information synthesis, their evaluation remains constrained by ad hoc empirical benchmarks. These heuristic approaches do not rigorously model agent behavior or adequately stress-test long-horizon synthesis and ambiguity resolution. To bridge this gap, we formalize DRA behavior through the lens of category theory, modeling deep research workflow as a composition of structure-preserving maps (functors). Grounded in this theoretical framework, we introduce a novel mechanism-aware benchmark with 296 questions designed to stress-test agents along four interpretable axes: traversing sequential connectivity chains, verifying intersections within V-structure pullbacks, imposing topological ordering on retrieved substructures, and performing ontological falsification via the Yoneda Probe. Our rigorous evaluation of 11 leading models establishes a persistently low baseline, with the state-of-the-art achieving only a 19.9\% average accuracy, exposing the difficulty of formal structural stress-testing. Furthermore, our findings reveal a stark dichotomy in the current AI capabilities. While advanced deep research pipelines successfully redefine dynamic topological re-ordering and exhibit robust ontological verification -- matching pure reasoning models in falsifying hallucinated premises -- they almost universally collapse on multi-hop structural synthesis. Crucially, massive performance variance across tasks exposes a lingering reliance on brittle heuristics rather than a systemic understanding. Ultimately, this work demonstrates that while top-tier autonomous agents can now organically unify search and reasoning, achieving a generalized mastery over complex structural information remains a formidable open challenge.\footnote{Our implementation will be available at https://github.com/tzq1999/CDR.

Motivation & Objective

  • Motivate the need for a rigorous, theory-grounded evaluation of DRAs beyond ad hoc benchmarks.
  • Introduce a category-theoretic formalization of DRA behavior and state spaces.
  • Propose a mechanism-aware benchmark to stress-test long-horizon synthesis and ambiguity resolution.
  • Quantify agent performance across multiple models to reveal structural strengths and weaknesses.

Proposed method

  • Model DRA behavior as a sequence of structure-preserving functors between categorical state spaces (Query, Web, Retrieved Subgraphs, and Reasoning).
  • Define precise category-theoretic notions (pullbacks, limits/colimits) to capture verification and aggregation tasks.
  • Design a 296-question benchmark organized along four axes: sequential connectivity, V-structure intersections, substructure ordering, and ontological falsification via Yoneda Probe.
  • Evaluate 11 leading models across reasoning, search-augmented, and autonomous DRA paradigms using a human-verified evaluation pipeline.

Experimental results

Research questions

  • RQ1Can category-theoretic abstractions faithfully model DRA search and reasoning workflows?
  • RQ2How well do current models preserve structural relationships (via functors) during search and reasoning tasks?
  • RQ3What are the primary failure modes of DRAs under long-horizon synthesis and ambiguity resolution?
  • RQ4Do DRAs exhibit strong ontological verification or do they rely on brittle heuristics across tasks?
  • RQ5How does performance vary across the four proposed categorical evaluation axes?

Key findings

  • State-of-the-art models achieve only 19.9% average accuracy on the benchmark.
  • Advanced DRA pipelines show strengths in dynamic topological reordering and ontological verification, comparable to pure reasoning models for falsifying hallucinated premises.
  • Models generally fail on multi-hop structural synthesis and exhibit blind spots under certain mathematical constraints.
  • There is large performance variance across tasks and models, indicating reliance on heuristics rather than systemic understanding.
  • The study highlights that achieving generalized mastery over complex structural information with DRAs remains a significant open challenge.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.