Skip to main content
QUICK REVIEW

[Paper Review] Process Mining for Unstructured Data: Challenges and Research Directions

Agnes Koschmider, Milda Aleknonytė‐Resch|arXiv (Cornell University)|Nov 30, 2023
Business Process Modeling and Analysis7 citations
TL;DR

This paper identifies key challenges in applying process mining to unstructured data—such as video, sensor streams, and audio—and proposes a five-step analytics pipeline to transform such data into process-mining-ready event logs. It outlines five critical research directions: hybrid modeling with domain knowledge, data fusion for comprehensive insights, advanced visualization, explainable AI integration, and ethical/legal frameworks, positioning process mining on unstructured data as a transformative yet underexplored frontier in process analytics.

ABSTRACT

The application of process mining for unstructured data might significantly elevate novel insights into disciplines where unstructured data is a common data format. To efficiently analyze unstructured data by process mining and to convey confidence into the analysis result, requires bridging multiple challenges. The purpose of this paper is to discuss these challenges, present initial solutions and describe future research directions. We hope that this article lays the foundations for future collaboration on this topic.

Motivation & Objective

  • To address the growing need for process mining in domains where unstructured data (e.g., video, sensor streams, audio) is prevalent but not natively compatible with traditional process mining.
  • To identify and systematize the core challenges in transforming low-level, unstructured data into structured event logs suitable for process mining.
  • To propose a systematic five-step analytics pipeline for preprocessing and analyzing unstructured data in process mining workflows.
  • To outline five forward-looking research directions that bridge technical, ethical, and domain-specific gaps in the field.
  • To lay the foundation for future collaboration by framing open challenges and actionable research agendas in unstructured data process mining.

Proposed method

  • Proposes a five-stage process analytics pipeline: data collection, event abstraction, event log creation, process mining, and result interpretation, tailored for unstructured data sources.
  • Employs data abstraction techniques to extract meaningful activities from raw unstructured inputs (e.g., video frames, sensor readings) and map them to symbolic event log entries.
  • Introduces hybrid modeling approaches that integrate domain knowledge into process mining workflows to improve model accuracy and relevance.
  • Advocates for data fusion techniques combining heterogeneous sources (e.g., video, sensors, logs) to create richer, more accurate process representations.
  • Promotes the use of interactive and scalable visualization methods—such as dashboards and 3D models—to convey complex process mining results to stakeholders.
  • Encourages integration of explainable AI (XAI) techniques to enhance transparency and trust in process mining outcomes derived from unstructured inputs.

Experimental results

Research questions

  • RQ1How can unstructured data such as video, audio, and sensor streams be systematically transformed into structured event logs suitable for process mining?
  • RQ2What are the key challenges in ensuring data quality and process fidelity when deriving event logs from low-level, unstructured data sources?
  • RQ3How can domain-specific knowledge be effectively integrated into process mining models to improve accuracy and relevance in real-world applications?
  • RQ4What data fusion strategies can combine heterogeneous data sources (e.g., video, sensor data, logs) to yield more comprehensive and reliable process insights?
  • RQ5How can ethical, legal, and privacy concerns—especially around surveillance data and bias—be proactively addressed in process mining of unstructured data?

Key findings

  • The transformation of unstructured data into process-mining-ready event logs requires multiple intermediary steps, including event abstraction and data structuring, due to the lack of inherent case IDs, activity names, and timestamps.
  • Traditional process mining assumes discrete, totally ordered, and accurate event logs; unstructured data often violates these assumptions, making direct application infeasible without preprocessing.
  • Data fusion of heterogeneous sources—such as video, sensor data, and production logs—can reveal patterns invisible in isolated data streams, enhancing process insight and monitoring capabilities.
  • Advanced visualization techniques, including interactive dashboards and 3D representations, are essential for conveying complex process mining results from unstructured data to non-technical stakeholders.
  • Explainability in process mining on unstructured data is critical for high-stakes domains like healthcare and law, and future systems must integrate feedback mechanisms to improve model trustworthiness.
  • Ethical and legal challenges—particularly around privacy, bias, and data protection—are significant and require dedicated frameworks, including anonymization techniques and transparent governance models, to ensure responsible use.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.