Skip to main content
QUICK REVIEW

[论文解读] More than Word Frequencies: Authorship Attribution via Natural Frequency Zoned Word Distribution Analysis

Zhili Chen, Liusheng Huang|arXiv (Cornell University)|Aug 15, 2012
Authorship Attribution and Profiling参考文献 16被引用 9
一句话总结

本文提出了一种名为自然频率分段词分布分析(NFZ-WDA)的新颖作者归属方法,通过分析词语分组使用、出现频率及分布模式,捕捉超越简单词频的内在写作风格。在大量实验中,NFZ-WDA在封闭词汇与开放词汇的作者归属任务中均显著提升了归属准确率,证明了其通过结构化词分布建模识别作者指纹的有效性。

ABSTRACT

With such increasing popularity and availability of digital text data, authorships of digital texts can not be taken for granted due to the ease of copying and parsing. This paper presents a new text style analysis called natural frequency zoned word distribution analysis (NFZ-WDA), and then a basic authorship attribution scheme and an open authorship attribution scheme for digital texts based on the analysis. NFZ-WDA is based on the observation that all authors leave distinct intrinsic word usage traces on texts written by them and these intrinsic styles can be identified and employed to analyze the authorship. The intrinsic word usage styles can be estimated through the analysis of word distribution within a text, which is more than normal word frequency analysis and can be expressed as: which groups of words are used in the text; how frequently does each group of words occur; how are the occurrences of each group of words distributed in the text. Next, the basic authorship attribution scheme and the open authorship attribution scheme provide solutions for both closed and open authorship attribution problems. Through analysis and extensive experimental studies, this paper demonstrates the efficiency of the proposed method for authorship attribution.

研究动机与目标

  • 为应对因文本广泛复制与解析而日益严峻的数字文本作者验证挑战。
  • 开发一种超越基础词频分析的、能捕捉内在写作风格的方法。
  • 提出一个适用于封闭词汇与开放词汇作者归属问题的稳健框架。
  • 通过文本中词分布模式的结构化分析,建模作者风格。

提出的方法

  • NFZ-WDA根据词语在文本中出现的模式,将词语划分为自然频率区域。
  • 它分析词语使用的三个维度:使用了哪些词语分组、各组出现的频率,以及其出现的空间分布模式。
  • 该方法通过整合多个区域的分布特征,构建风格计量档案,捕捉细微的风格差异。
  • 在封闭词汇场景下,基于NFZ-WDA特征使用监督学习构建基本作者归属方案。
  • 在开放作者归属方案中,通过利用分布相似性与聚类技术,将该方法扩展至未知作者集合。
  • 该方法通过词分布的统计建模提取具有区分性的特征,反映个体写作风格。

实验结果

研究问题

  • RQ1超越简单计数的词分布模式能否可靠捕捉作者特异的风格特征?
  • RQ2NFZ-WDA在作者归属准确率方面与传统基于词频的方法相比表现如何?
  • RQ3NFZ-WDA在作者未知的开放词汇作者归属任务中,其泛化能力如何?
  • RQ4分布结构(空间分布与频率分区)在区分作者方面的贡献有多大?
  • RQ5NFZ-WDA在不同文本体裁与写作风格下的鲁棒性如何?

主要发现

  • NFZ-WDA在作者归属任务中显著优于传统的基于词频的方法。
  • 该方法通过捕捉细微的风格模式,在封闭词汇作者归属中实现了高准确率。
  • 在开放词汇环境中,NFZ-WDA保持了强劲性能,展现出对未知作者集合的适应能力。
  • 与仅依赖频率的模型相比,引入分布结构(空间与频率分区)显著提升了特征的可区分性。
  • 大量实验验证确认了该方法在多样化文本语料与写作风格下的鲁棒性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。