Skip to main content
QUICK REVIEW

[论文解读] Computational analyses of the topics, sentiments, literariness, creativity and beauty of texts in a large Corpus of English Literature

Arthur M. Jacobs, Annette Kinder|arXiv (Cornell University)|Jan 12, 2022
Sentiment Analysis and Opinion Mining被引用 8
一句话总结

本研究运用计算方法对古腾堡文学英语语料库(GLEC)进行分析,探讨六种文学类别及逾百位作者在主题、情感、文学性、创造力和感知美感方面的特征。研究提出新颖的语义复杂性度量指标——文本内差异性和逐步距离,识别出戏剧为最具文学性与创造力的体裁,诗歌与戏剧在创造力方面得分最高,而《爱玛》被预测为简·奥斯汀小说中最美的一部,这些特征在文本分类与作者识别任务中的准确率可达75%至97%。

ABSTRACT

The Gutenberg Literary English Corpus (GLEC, Jacobs, 2018a) provides a rich source of textual data for research in digital humanities, computational linguistics or neurocognitive poetics. In this study we address differences among the different literature categories in GLEC, as well as differences between authors. We report the results of three studies providing i) topic and sentiment analyses for six text categories of GLEC (i.e., children and youth, essays, novels, plays, poems, stories) and its >100 authors, ii) novel measures of semantic complexity as indices of the literariness, creativity and book beauty of the works in GLEC (e.g., Jane Austen's six novels), and iii) two experiments on text classification and authorship recognition using novel features of semantic complexity. The data on two novel measures estimating a text's literariness, intratextual variance and stepwise distance (van Cranenburgh et al., 2019) revealed that plays are the most literary texts in GLEC, followed by poems and novels. Computation of a novel index of text creativity (Gray et al., 2016) revealed poems and plays as the most creative categories with the most creative authors all being poets (Milton, Pope, Keats, Byron, or Wordsworth). We also computed a novel index of perceived beauty of verbal art (Kintsch, 2012) for the works in GLEC and predict that Emma is the theoretically most beautiful of Austen's novels. Finally, we demonstrate that these novel measures of semantic complexity are important features for text classification and authorship recognition with overall predictive accuracies in the range of .75 to .97. Our data pave the way for future computational and empirical studies of literature or experiments in reading psychology and offer multiple baselines and benchmarks for analysing and validating other book corpora.

研究动机与目标

  • 探究在大型英语文学语料库中,主要文学类别在主题、情感与风格特征方面的差异。
  • 开发并验证文学文本中文学性、创造力与感知美感的新型计算度量方法。
  • 评估语义复杂性特征在文本分类与作者识别任务中的实用性。
  • 为未来数字人文、神经认知诗学与计算语言学研究提供实证基准与基线数据。

提出的方法

  • 对六种文学类别(儿童与青少年读物、随笔、小说、戏剧、诗歌与故事)进行了主题与情感分析。
  • 计算了新颖的语义复杂性指标——文本内差异性与逐步距离,以评估文学性。
  • 应用基于Gray等人(2016)的创造力指数,评估不同作者与体裁的语言原创性。
  • 采用基于Kintsch(2012)的感知美感指数,估算文学作品中的语言艺术性。
  • 利用语义复杂性特征训练文本分类与作者识别模型,并通过准确率指标评估性能。
  • 所有分析均在古腾堡文学英语语料库(GLEC)上完成,涵盖逾百位作者与多种文学形式。

实验结果

研究问题

  • RQ1在GLEC语料库中,不同文学体裁(如戏剧、诗歌、小说)在主题分布与情感特征上存在何种差异?
  • RQ2文本内差异性与逐步距离在多大程度上可作为文本文学性可靠指标?
  • RQ3根据Gray等人(2016)的指数,哪些文学体裁与作者展现出最高的语言创造力?
  • RQ4语义复杂性特征能否显著提升文本分类与作者识别模型的准确率?
  • RQ5基于Kintsch(2012)的感知美感指数,简·奥斯汀的哪部小说被预测为最美丽?

主要发现

  • 基于文本内差异性与逐步距离指标,戏剧被识别为最具文学性的体裁,其次为诗歌与小说。
  • 诗歌与戏剧在创造力方面表现最突出,所有创造力最高的个体作者均为诗人,如弥尔顿、蒲柏、济慈、拜伦与华兹华斯。
  • 根据Kintsch(2012)的感知美感指数,小说《爱玛》被预测为简·奥斯汀作品中理论上最美的一部。
  • 语义复杂性特征在文本分类与作者识别任务中实现了75%至97%的预测准确率。
  • 本研究为文学文本的文学性、创造力与美感提供了经验证的计算基准,为未来数字人文与认知科学的研究奠定了基础。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。