[논문 리뷰] Talk Freely, Execute Strictly: Schema-Gated Agentic AI for Flexible and Reproducible Scientific Workflows
이 논문은 대화 의도를 실행과 분리하기 위한 schema-gated orchestration을 도입하고, 과학 워크플로우에서 에이전트 AI를 분석하는 20개 시스템을 ED/CF 축으로 분석하며, 유연성과 결정적 실행을 모두 달성하기 위한 참조 아키텍처를 제안한다.
Large language models (LLMs) can now translate a researcher's plain-language goal into executable computation, yet scientific workflows demand determinism, provenance, and governance that are difficult to guarantee when an LLM decides what runs. Semi-structured interviews with 18 experts across 10 industrial R&D stakeholders surface 2 competing requirements--deterministic, constrained execution and conversational flexibility without workflow rigidity--together with boundary properties (human-in-the-loop control and transparency) that any resolution must satisfy. We propose schema-gated orchestration as the resolving principle: the schema becomes a mandatory execution boundary at the composed-workflow level, so that nothing runs unless the complete action--including cross-step dependencies--validates against a machine-checkable specification. We operationalize the 2 requirements as execution determinism (ED) and conversational flexibility (CF), and use these axes to review 20 systems spanning 5 architectural groups along a validation-scope spectrum. Scores are assigned via a multi-model protocol--15 independent sessions across 3 LLM families--yielding substantial-to-near-perfect inter-model agreement (Krippendorff a=0.80 for ED and a=0.98 for CF), demonstrating that multi-model LLM scoring can serve as a reusable alternative to human expert panels for architectural assessment. The resulting landscape reveals an empirical Pareto front--no reviewed system achieves both high flexibility and high determinism--but a convergence zone emerges between the generative and workflow-centric extremes. We argue that a schema-gated architecture, separating conversational from execution authority, is positioned to decouple this trade-off, and distill 3 operational principles--clarification-before-execution, constrained plan-act orchestration, and tool-to-workflow-level gating--to guide adoption.
연구 동기 및 목표
- AI-주도 과학 워크플로우에서 실행 결정성과 대화 가능성의 유연성 간의 균형을 정하는 실무자 요구사항 식별.
- 현존 시스템을 실행 결정성(ED)과 대화 가능성(CF) 설계 공간에 매핑한다.
- LLM 계열 전반에 걸친 아키텍처 평가를 위한 모델 간 점수 산정 신뢰성을 입증한다.
- ED/CF 트레이드오프에 대한 원칙 있는 해결책으로 schema-gated orchestration을 제안한다.
- 실제 워크플로우에서의 채택을 안내하기 위한 참조 아키텍처와 세 가지 작동 원칙을 제시한다.]
- method:["반구조화된 인터뷰를 10개 산업 연구개발 이해관계자에서 18명의 전문가를 대상으로 수행하여 요구사항과 경계 속성을 도출한다.","다섯 가지 아키텍처 그룹에 걸친 20개의 대표 시스템을 검토하고 ED 및 CF 축에서 다섯점 등급 규칙을 사용해 점수를 매긴다.","세 가지 LLM 계열(ChatGPT, Claude, Gemini)에서 15회의 독립 점수 매김 세션을 수행하여 모델 간 일치도(Krippendorff’s α)를 평가한다.","설계 공간을 분석하여 경험적 Pareto 프런트를 드러내고 패러다임 간 수렴 영역을 식별한다.","schema-gated orchestration을 설계 원칙으로 삼고 세 가지 작동 신조를 제시하며 원천 보장을 포함한 참조 아키텍처를 개략적으로 제시한다."]
- research_questions:["AI-주도 과학 워크플로우에서 실행 결정성과 대화 가능성의 유연성 두가지를 달성하기 위해 필요한 아키텍처적 요구사항은 무엇인가?","현 시스템은 ED/CF에 어떻게 정렬되며, 패러다임(생성형, 도구 보강, 스키마-게이티드, 워크플로우 기반) 간 어떤 트레이드오프가 존재하는가?","schema-gated orchestration이 대화 권한과 실행 권한을 분리하여 재현성과 거버넌스를 향상시킬 수 있는가?","구성된 워크플로우 전반에 걸친 schema-gated 실행을 구현하기 위한 실용적 함의와 아키텍처 패턴은 무엇인가?"],"key_findings:["경험적 트레이드오프가 존재한다: 검토된 어떤 시스템도 높은 유연성과 높은 실행 결정성 두 가지를 모두 달성하지 못한다(Pareto front).","모델 간 점수 산정 신뢰성은 상당히-거의 완벽에 가까운 수준이다(Krippendorff’s α = 0.80 for ED and 0.98 for CF) across 15 scoring runs over three LLM families.","Schema-gated orchestration은 개별 도구 호출의 스키마 검증을 구성 워크플로우 계획으로 확장하여 결정적 실행과 대화 가능성을 더 잘 지원할 수 있다.","두 개의 작동 영역이 도출된다: 이상적에 가장 근접한 schema-gated 그룹(ID 8–9), 워크플로우 중심 및 워크플로우+NL 그룹이 수렴하여 ED가 높아지지만 CF가 낮아지는 경향.","세 가지 작동 원칙이 제시된다: 실행 전 명확화(clarification-before-execution), 제약된 계획–실행 오케스트레이션, 도구에서 워크플로우 수준으로의 게이팅.","참조 아키텍처가 제안되어 스키마 검증된 레지스트리와 대화 계층을 오케스트레이션 컨트롤러를 통해 분리하고 엔드투엔드 출처 보장을 가능하게 한다."]
- table_headers:
- table_rows:
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.