Skip to main content
QUICK REVIEW

[Paper Review] Early Evidence of Vibe-Proving with Consumer LLMs: A Case Study on Spectral Region Characterization with ChatGPT-5.2 (Thinking)

Brecht Verbeken, Brando Vagenende|arXiv (Cornell University)|Feb 21, 2026
Artificial Intelligence in Healthcare and Education0 citations
TL;DR

The paper presents a structured, auditable case study where a consumer LLM (ChatGPT-5.2 Thinking) collaborates with humans to prove a spectral-region characterization for a 4-cycle row-stochastic matrix family, highlighting workflow, verification bottlenecks, and the potential of human-in-the-loop theorem proving.

ABSTRACT

Large Language Models (LLMs) are increasingly used as scientific copilots, but evidence on their role in research-level mathematics remains limited, especially for workflows accessible to individual researchers. We present early evidence for vibe-proving with a consumer subscription LLM through an auditable case study that resolves Conjecture 20 of Ran and Teng (2024) on the exact nonreal spectral region of a 4-cycle row-stochastic nonnegative matrix family. We analyze seven shareable ChatGPT-5.2 (Thinking) threads and four versioned proof drafts, documenting an iterative pipeline of generate, referee, and repair. The model is most useful for high-level proof search, while human experts remain essential for correctness-critical closure. The final theorem provides necessary and sufficient region conditions and explicit boundary attainment constructions. Beyond the mathematical result, we contribute a process-level characterization of where LLM assistance materially helps and where verification bottlenecks persist, with implications for evaluation of AI-assisted research workflows and for designing human-in-the-loop theorem proving systems.

Motivation & Objective

  • Demonstrate that consumer LLMs can contribute to mathematically substantive proof development with explicit human verification.
  • Provide an auditable artifact set (transcripts and proof drafts) for end-to-end inspection of AI-assisted theorem proving.
  • Characterize the division of labor between LLM-generated structure and human correctness-critical verification.

Proposed method

  • Adopt a generate–referee–repair workflow with seven ChatGPT-5.2 (Thinking) threads and four versioned drafts.
  • Apply a Dmitriev–Dynkin trigonometric reduction to the 4-cycle row-stochastic matrix family.
  • Use a target theorem statement and boundary information as scaffolding for LLM-proposed proof strategies.
  • Incorporate explicit correctness obligations (quadrant handling, endpoint admissibility, and algebraic expansions) and patch-search against independent sessions.
  • Leverage Lamport-style claim decomposition to organize dependencies and verification steps.

Experimental results

Research questions

  • RQ1Can a consumer-access LLM contribute to a research-level mathematical proof with auditable, end-to-end workflow?
  • RQ2What is the division of labor between AI-generated structure and human verification in vibe-proving for spectral-region problems?
  • RQ3What verification bottlenecks arise, and how can workflow practices mitigate them?
  • RQ4Is it feasible to produce a complete, checkable characterization of non-real eigenvalues for a 4-cycle row-stochastic matrix family using LLM-assisted proof?
  • RQ5How do transcript-based artefacts and versioning support auditability in AI-assisted mathematics?

Key findings

  • A stable generate–referee–repair loop yields a complete and checkable proof for the conjecture’s spectral-region characterization.
  • LLMs are strongest at proposing global structure and algebraic shortcuts, while humans handle correctness-critical verification and long expansions.
  • The verification bottleneck concentrates on a few key obligations (e.g., tight-regime inequalities and factorization steps) that are amenable to mechanized checking.
  • Parallel patch search, bounded referee passes, and version-controlled rewriting reduce regressions and improve auditability.
  • An auditable workflow with explicit artefacts can illuminate where AI-assisted workflows help and where human verification remains essential.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.