[Paper Review] Effect of a Process Mining based Pre-processing Step in Prediction of the Critical Health Outcomes
This study proposes a process mining-based pre-processing step using the concatenation algorithm to reduce complexity in healthcare event logs, improving process model quality and prediction accuracy for critical outcomes like mortality and readmission. Applied to 16 MIMIC-III and University of Illinois Hospital datasets, the method enhanced F-Measure (up to 0.95), fitness, and AUC in DREAM-based predictions, demonstrating improved model fidelity and predictive performance.
Predicting critical health outcomes such as patient mortality and hospital readmission is essential for improving survivability. However, healthcare datasets have many concurrences that create complexities, leading to poor predictions. Consequently, pre-processing the data is crucial to improve its quality. In this study, we use an existing pre-processing algorithm, concatenation, to improve data quality by decreasing the complexity of datasets. Sixteen healthcare datasets were extracted from two databases - MIMIC III and University of Illinois Hospital - converted to the event logs, they were then fed into the concatenation algorithm. The pre-processed event logs were then fed to the Split Miner (SM) algorithm to produce a process model. Process model quality was evaluated before and after concatenation using the following metrics: fitness, precision, F-Measure, and complexity. The pre-processed event logs were also used as inputs to the Decay Replay Mining (DREAM) algorithm to predict critical outcomes. We compared predicted results before and after applying the concatenation algorithm using Area Under the Curve (AUC) and Confidence Intervals (CI). Results indicated that the concatenation algorithm improved the quality of the process models and predictions of the critical health outcomes.
Motivation & Objective
- To address the challenge of high data complexity in healthcare event logs that hinders accurate prediction of critical outcomes.
- To evaluate whether pre-processing with the concatenation algorithm improves the quality of process models derived from electronic health record data.
- To assess the impact of pre-processing on the predictive performance of DREAM-based models for patient mortality and hospital readmission.
- To quantify improvements in process model metrics—fitness, precision, F-Measure, and complexity—before and after concatenation.
- To demonstrate the value of integrating process mining techniques into clinical prediction pipelines.
Proposed method
- Sixteen healthcare datasets were extracted from MIMIC-III and University of Illinois Hospital databases and converted into event logs.
- The concatenation algorithm was applied to pre-process event logs, reducing complexity by merging similar sequences and minimizing concurrency.
- Split Miner (SM) was used to generate process models from both raw and pre-processed event logs for comparison.
- Model quality was evaluated using fitness, precision, F-Measure, and complexity metrics to assess structural fidelity.
- The pre-processed event logs were input into the Decay Replay Mining (DREAM) algorithm to predict critical health outcomes.
- Prediction performance was measured using Area Under the Curve (AUC) and 95% Confidence Intervals (CI), comparing results before and after pre-processing.
Experimental results
Research questions
- RQ1Does applying the concatenation pre-processing algorithm improve the quality of process models derived from healthcare event logs?
- RQ2To what extent does pre-processing reduce data complexity and enhance fitness, precision, and F-Measure in process models?
- RQ3How does concatenation-based pre-processing affect the predictive performance of DREAM in forecasting patient mortality and readmission?
- RQ4Are the improvements in model quality and prediction accuracy statistically significant, as indicated by AUC and confidence intervals?
- RQ5Can process mining techniques be effectively leveraged to enhance clinical prediction systems through data pre-processing?
Key findings
- The concatenation algorithm significantly improved process model quality, with F-Measure increasing to 0.95 and fitness reaching 0.95 in the best-performing case.
- Precision improved to 0.85 after pre-processing, indicating better alignment between the model and observed behavior.
- The complexity of the process models was reduced, suggesting a more concise and interpretable representation of clinical workflows.
- DREAM-based predictions showed improved AUC after pre-processing, indicating higher discriminative ability in identifying critical health outcomes.
- The 95% Confidence Intervals for AUC values confirmed the statistical robustness of the performance gains across multiple datasets.
- Overall, the integration of concatenation pre-processing led to more accurate and reliable predictions of patient mortality and hospital readmission.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.