Skip to main content
QUICK REVIEW

[Paper Review] Conversational Process Modeling: Can Generative AI Empower Domain Experts in Creating and Redesigning Process Models?

Nataliia Klievtsova, Janik-Vasily Benzin|arXiv (Cornell University)|Apr 19, 2023
Business Process Modeling and AnalysisBusiness, Management and Accounting3 citations
TL;DR

This paper investigates how generative AI chatbots can assist domain experts in creating and redesigning business process models through conversational modeling. It proposes a framework for evaluating LLM-generated models using KPIs, prompts, and user surveys, finding that while chatbots can produce reasonably complete and correct models (43% accuracy in control flow), human oversight remains essential due to variability in quality and semantic understanding limitations.

ABSTRACT

AI-driven chatbots such as ChatGPT have caused a tremendous hype lately. For BPM applications, several applications for AI-driven chatbots have been identified to be promising to generate business value, including explanation of process mining outcomes and preparation of input data. However, a systematic analysis of chatbots for their support of conversational process modeling as a process-oriented capability is missing. This work aims at closing this gap by providing a systematic analysis of existing chatbots. Application scenarios are identified along the process life cycle. Then a systematic literature review on conversational process modeling is performed, resulting in a taxonomy of application scenarios for conversational process modeling, including paraphrasing and improvement of process descriptions. In addition, this work suggests and applies an evaluation method for the output of AI-driven chatbots with respect to completeness and correctness of the process models. This method consists of a set of KPIs on a test set, a set of prompts for task and control flow extraction, as well as a survey with users. Based on the literature and the evaluation, recommendations for the usage (practical implications) and further development (research directions) of conversational process modeling are derived.

Motivation & Objective

  • Address the gap in systematic analysis of chatbots for conversational process modeling (ConverMod) in business process management (BPM).
  • Identify and categorize application scenarios for ConverMod across the business process life cycle.
  • Develop and apply an evaluation framework to assess LLM-generated process models for completeness and correctness.
  • Provide practical implications and research directions for integrating chatbots into BPM workflows.
  • Investigate the role of prompt engineering, textual representation, and user expertise in model quality outcomes.

Proposed method

  • Conduct a systematic literature review to develop a taxonomy of ConverMod application scenarios, including paraphrasing and improving process descriptions.
  • Design a test set based on higher education process data and use it to evaluate LLMs (GPT-3.5, GPT-4) on task and control flow extraction.
  • Generate process models using Mermaid.js and Graphviz from LLM outputs to enable visual validation.
  • Define a set of KPIs to quantitatively assess model completeness and correctness on the test set.
  • Perform a qualitative evaluation via a user survey comparing gold-standard models with LLM-generated ones across completeness and correctness.
  • Integrate findings into a methodological framework for evaluating ConverMod capabilities, applicable to both proprietary and open-source LLMs.

Experimental results

Research questions

  • RQ1RQ1: How can conversational modeling methods/tools be employed for process modeling?
  • RQ2RQ2: Which conversational modeling methods/tools exist for process modeling?
  • RQ3RQ3: How can conversational modeling methods/tools be evaluated with respect to process modeling quality?
  • RQ4RQ4: What are the practical and research implications of chatbots for BPM modeling practice and research?

Key findings

  • LLM-generated process models achieved an average of 43% correctness in control flow evaluation, indicating significant room for improvement.
  • Despite low accuracy in semantic correctness, users preferred LLM-generated models over gold-standard models in surveys, regardless of their process modeling experience.
  • Model completeness was higher in LLM outputs, but correctness—especially in control flow and semantic structure—remained a major challenge.
  • The quality of generated models was highly sensitive to prompt design, textual representation, and the complexity of the input description.
  • Domain experts face difficulty selecting the best model from multiple LLM outputs, even when a gold-standard model exists.
  • Current off-the-shelf chatbots are not yet reliable for advanced tasks like model comparison, querying, or refactoring due to poor understanding of process model semantics.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.