Skip to main content
QUICK REVIEW

[论文解读] Open ERP System Data For Occupational Fraud Detection

Julian Tritscher, Fabian Gwinner|arXiv (Cornell University)|Jun 9, 2022
Imbalanced Data Classification Techniques被引用 4
一句话总结

本文提出了一种新颖的方法,用于生成包含正常业务操作和真实职业欺诈场景的公开合成ERP系统数据。通过利用严肃游戏模拟真实ERP系统中的用户交互,作者生成了具有细粒度欺诈标注的多年期数据集,从而实现企业资源计划环境中欺诈检测技术的可重现基准测试。

ABSTRACT

Recent estimates report that companies lose 5% of their revenue to occupational fraud. Since most medium-sized and large companies employ Enterprise Resource Planning (ERP) systems to track vast amounts of information regarding their business process, researchers have in the past shown interest in automatically detecting fraud through ERP system data. Current research in this area, however, is hindered by the fact that ERP system data is not publicly available for the development and comparison of fraud detection methods. We therefore endeavour to generate public ERP system data that includes both normal business operation and fraud. We propose a strategy for generating ERP system data through a serious game, model a variety of fraud scenarios in cooperation with auditing experts, and generate data from a simulated make-to-stock production company with multiple research participants. We aggregate the generated data into ready to used datasets for fraud detection in ERP systems, and supply both the raw and aggregated data to the general public to allow for open development and comparison of fraud detection approaches on ERP system data.

研究动机与目标

  • 为解决职业欺诈检测研究中缺乏公开可用的ERP系统数据的问题。
  • 生成包含正常操作和多样化欺诈场景的真实合成ERP数据。
  • 通过发布原始数据和聚合数据集,确保欺诈检测方法的可重现性和可比性。
  • 将数据生成范围从P2P流程扩展至O2C流程,以提升适用范围。
  • 在特征层面提供细粒度的专家标注欺诈标签,以支持对检测模型性能的深入评估。

提出的方法

  • 利用严肃游戏模拟真实ERP系统中的用户交互,捕获正常和欺诈性业务流程。
  • 与审计专家合作建模多种职业欺诈场景,以确保真实性和多样性。
  • 在多个财年中进行多次模拟运行,以生成纵向ERP数据。
  • 从ERP系统中提取原始交易数据,并将其聚合为即用型数据集。
  • 通过连接财务会计表创建联合数据集,排除非区分性欺诈案例以保持数据完整性。
  • 执行专家级别的特征标注,突出显示可能暗示潜在欺诈案例的异常列条目。

实验结果

研究问题

  • RQ1严肃游戏能否有效用于生成包含正常操作和职业欺诈的、真实且公开的ERP数据?
  • RQ2如何对合成ERP数据进行结构化设计,以支持欺诈检测算法的可重现基准测试?
  • RQ3为实现对检测性能的有意义评估,欺诈标注的粒度应达到何种程度?
  • RQ4当数据在财务会计层面聚合时,ERP数据中的欺诈案例在多大程度上仍能与正常交易区分开来?
  • RQ5如何将数据生成策略扩展至涵盖O2C和P2P等多个业务流程,以构建统一的开放数据集?

主要发现

  • 作者成功通过基于严肃游戏的模拟,生成了包含正常和欺诈交易的多年期ERP系统数据。
  • 该数据集涵盖多个财年中的86起欺诈案例,其中50起在最终联合数据集中被标注。
  • 发现欺诈案例主要通过少量异常条目在看似正常的交易中被识别,支持使用统计异常检测方法。
  • 专家标注的细粒度特征级标签使能够在交易层面进行欺诈检测性能的深入分析。
  • 最终数据集已公开发布,包含原始数据和聚合数据,以支持ERP欺诈检测领域的开放、可重现研究。
  • 该方法可实现对不同业务流程(包括O2C和P2P)中欺诈检测方法的基准测试,并支持真实欺诈场景。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。