[论文解读] Towards A Better Understanding Of The Leading Digits Phenomena
本文通过区分随机过程与确定性过程在首位数分布中的表现,推进了对本福德定律的理解。它识别出解释对数行为的密度曲线中共享的内在特征,提出一种新的法医学欺诈检测方法——即使在数据已知本福德定律的情况下也有效,并揭示了随机系统与确定性系统在数字符合规律性方面存在根本差异,挑战了以往关于求和相等性的假设。
That the logarithmic distribution manifests itself in the random as well as in the deterministic (multiplication processes) has long intrigued researchers in Benford's Law. In this article it is argued that it springs from one common intrinsic feature of their density curves. On the other hand, the profound dichotomy between the random and the deterministic in the context of Benford's Law is noted here, acknowledging the need to distinguish between them. From its very inception, the field has been suffering from a profound confusion and mixing of these two very different logarithmic flavors, causing mistaken conclusions. One example is Allaart's proof of equality of sums along digital lines, which can only be applied to deterministic processes. Random data lack this equality and consistently show significantly larger sums for lower digits, thus rendering any attempt at test of summation equality irrelevant and futile in the context of forensic analysis regarding accounting and financial fraud detection. Another digital regularity is suggested here, one that is found in logarithmic as well as non-logarithmic random data sets. In addition, chains of distributions that are linked via parameter selection are found to be logarithmic, either in the limit where the number of the sequences in the chain approaches infinity, or where the distributions generating the parameters are themselves logarithmic. A new forensic data analysis method in the context of fraud detection is suggested here even for data types that do not obey Benford's Law, and in particularly regarding tax evasion applications. This can also serve as a robust forensic tool to investigate fraudulent fake data provided by the sophisticated cheater already aware of Benford's Law, a challenge that would become increasing problematic to tax authorities in the future as Benford's Law becomes almost common knowledge.
研究动机与目标
- 澄清随机过程与确定性过程在本福德定律现象中的根本区别。
- 解决因混淆随机与确定性对数行为而长期困扰文献的困惑。
- 提出一种即使在数据不遵循本福德定律时也有效的新型法医学数据分析方法。
- 识别出在对数与非对数随机数据集中均有效的新型数字符合规律性。
- 证明参数依赖分布链在特定条件下收敛于对数形式。
提出的方法
- 分析随机与确定性过程在产生对数首位数分布时共有的密度曲线内在几何特征。
- 对比随机数据与确定性数据在数字线上求和行为的差异,表明仅在确定性情况下求和相等成立。
- 提出一种基于非本福德数据中可观测数字符合规律性的新型法医学测试,适用于检测复杂欺诈行为。
- 研究通过参数选择关联的分布链,证明其在极限情况下或当生成分布为对数形式时,会趋向对数形式。
- 通过密度曲线与分布序列的数学分析,推导出对数行为出现的条件。
- 将研究成果应用于开发一种针对税务欺诈与金融数据分析的稳健欺诈检测工具。
实验结果
研究问题
- RQ1在随机与确定性过程的密度曲线中,何种共同特征可解释对数首位数分布的出现?
- RQ2为何以往在数字线上应用求和相等性的尝试在随机数据中失败?这对法医学分析有何影响?
- RQ3能否为不遵循本福德定律的数据开发出可靠的法医学方法,特别是当欺诈者已知晓本福德定律时?
- RQ4参数依赖分布链在何种条件下收敛于对数形式?
- RQ5在随机数据中存在何种新型数字符合规律性,且不依赖于本福德定律?
主要发现
- 本福德定律中的对数分布源于随机与确定性过程在密度曲线中共享的内在特征。
- 沿数字线的求和相等性,此前被普遍假设成立,实际上仅在确定性过程中成立,而在随机数据中不成立,因此此类测试在随机情境下无法用于法医学分析。
- 随机数据始终显示较低数字的求和更大,与相等性假设相矛盾,使其无法作为诊断工具使用。
- 提出一种新型法医学方法,适用于非本福德数据,为应对知晓本福德定律的复杂欺诈者提供了稳健工具。
- 通过参数选择关联的分布链在极限情况下或当生成分布为对数形式时,会收敛为对数形式。
- 识别出一种新型数字符合规律性,其存在于对数与非对数随机数据集中,且不依赖于本福德定律。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。