[论文解读] Detection of statistically significant differences between process variants through declarative rules
本文提出声明规则变体分析(DRVA),一种通过自然语言表达的声明规则检测过程变体之间统计显著差异的方法。通过结合Declare约束与统计显著性检验,DRVA在保证统计严谨性的前提下,识别出整个过程运行中的高层次行为差异,在执行时间和可解释性方面优于现有最先进方法。
Services and products are often offered via the execution of processes that vary according to the context, requirements, or customisation needs. The analysis of such process variants can highlight differences in the service outcome or quality, leading to process adjustments and improvement. Research in the area of process mining has provided several methods for process variants analysis. However, very few of those account for a statistical significance analysis of their output. Moreover, those techniques detect differences at the level of process traces, single activities, or performance. In this paper, we aim at describing the distinctive behavioural characteristics between variants expressed in the form of declarative process rules. The contribution to the research area is two-pronged: the use of declarative rules for the explanation of the process variants and the statistical significance analysis of the outcome. We assess the proposed method by comparing its results to the most recent process variants analysis methods. Our results demonstrate not only that declarative rules reveal differences at an unprecedented level of expressiveness, but also that our method outperforms the state of the art in terms of execution time.
研究动机与目标
- 解决现有过程变体分析技术中缺乏统计显著性评估的问题。
- 通过以声明规则而非原始过程图或追踪级别特征表达差异,提升结果的可解释性。
- 利用自然语言规则实现可扩展的、高层次的行为比较,对过程变体进行全局分析。
- 通过应用统计显著性检验,确保检测到的差异并非由随机波动引起。
- 在执行效率和解释能力两方面均超越当前最先进方法。
提出的方法
- 该方法使用Declare约束(如Response、Precedence和AlternateResponse)作为形式化构建模块,以表示行为规则。
- 通过比较不同变体中活动的频率和共现模式,推断出与变体相关的特定规则。
- 采用基于置换的检验方法评估统计显著性,以判断观察到的规则差异在原假设(即变体间无差异)下是否极不可能发生。
- 输出结果为按重要性排序的声明规则列表,这些规则在语义上具有区分性且经统计显著性验证,以自然语言表达,便于人类阅读。
- 该方法作用于完整的过程运行,而非局部事件对,从而实现全局、上下文感知的规则推断。
- 若某一规则在一个变体中成立而在另一变体中不成立,则认为其具有区分性;其显著性通过重采样方法估算p值进行验证。
实验结果
研究问题
- RQ1声明规则能否以人类可读的方式有效表达过程变体之间的高层次行为差异?
- RQ2与不考虑显著性的方法相比,引入统计显著性检验是否能提高检测差异的可靠性?
- RQ3所提出方法在执行时间与可扩展性方面与最先进技术相比表现如何?
- RQ4该方法能否有效处理具有高追踪多样性、结构灵活的复杂过程变体,如BPIC15日志中的情况?
- RQ5声明规则在表达力与语义意义方面,相较于基于图或追踪级别的分析,优势有多大?
主要发现
- 在BPIC13、BPIC15和SEPSIS日志上,DRVA的执行时间分别为4.9秒、326.7秒和4.6秒,在所有情况下均优于MFVA和PESVA。
- 在BPIC15日志中,PESVA在3小时后超时,而DRVA仅用326.7秒完成,展现出更优的可扩展性。
- DRVA通过以声明规则表达差异,而非复杂图谱或模糊的事件级陈述,相比MFVA和PESVA产生了更具可解释性的结果。
- 在RTFMP日志中,DRVA识别出10条具有区分性的规则,例如在变体B中‘若16_LGSD_010发生,则01_HOOFD_490_2必须发生’,并以统计置信度确认。
- 在SEPSIS案例中,DRVA正确捕捉到变体A中‘01_HOOFD_492_2最多可发生一次’的特定规则,而PESVA未能清晰表达该规则。
- DRVA采用全局、基于规则的推理方式,相比依赖局部事件对或主要事件结构的方法,产生了更一致且语义更清晰的差异区分。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。