[論文レビュー] Computational analyses of the topics, sentiments, literariness, creativity and beauty of texts in a large Corpus of English Literature
本研究では、Gutenberg Literary English Corpus (GLEC) を用いて、6つの文学的ジャンルおよび100人以上の著者を対象に、トピック、感情、文学的特徴、創造性、および美的感受性を計算的手法で分析した。本研究では、文章内ばらつきと段階的距離という、新規の意味的複雑性指標を導入し、演劇が最も文学的で創造的であることを特定した。詩と演劇は創造性において最高得点を記録し、エマはオーステンの小説の中で最も美的であると予測された。これらの特徴は、テキスト分類および著者識別タスクにおいて75–97%の正確性を達成した。
The Gutenberg Literary English Corpus (GLEC, Jacobs, 2018a) provides a rich source of textual data for research in digital humanities, computational linguistics or neurocognitive poetics. In this study we address differences among the different literature categories in GLEC, as well as differences between authors. We report the results of three studies providing i) topic and sentiment analyses for six text categories of GLEC (i.e., children and youth, essays, novels, plays, poems, stories) and its >100 authors, ii) novel measures of semantic complexity as indices of the literariness, creativity and book beauty of the works in GLEC (e.g., Jane Austen's six novels), and iii) two experiments on text classification and authorship recognition using novel features of semantic complexity. The data on two novel measures estimating a text's literariness, intratextual variance and stepwise distance (van Cranenburgh et al., 2019) revealed that plays are the most literary texts in GLEC, followed by poems and novels. Computation of a novel index of text creativity (Gray et al., 2016) revealed poems and plays as the most creative categories with the most creative authors all being poets (Milton, Pope, Keats, Byron, or Wordsworth). We also computed a novel index of perceived beauty of verbal art (Kintsch, 2012) for the works in GLEC and predict that Emma is the theoretically most beautiful of Austen's novels. Finally, we demonstrate that these novel measures of semantic complexity are important features for text classification and authorship recognition with overall predictive accuracies in the range of .75 to .97. Our data pave the way for future computational and empirical studies of literature or experiments in reading psychology and offer multiple baselines and benchmarks for analysing and validating other book corpora.
研究の動機と目的
- 大規模な英語文学コーパスにおいて、主な文学的ジャンルごとのトピック、感情、スタイル的特徴の違いを調査すること。
- 文学的テキストにおける文学的特徴、創造性、美的感受性の新たな計算的測定法を構築し、妥当性を検証すること。
- 意味的複雑性特徴がテキスト分類および著者識別タスクにどの程度有用であるかを評価すること。
- 今後のデジタル・ヒューマニティーズ、神経認知的ポエティクス、計算言語学研究のための実証的ベンチマークとベースラインを提供すること。
提案手法
- 6つの文学的ジャンル(児童・若年層向け、随筆、小説、演劇、詩、物語)を対象に、トピック分析および感情分析を実施した。
- 文学的特徴の評価に、文章内ばらつきと段階的距離という、新規の意味的複雑性インデックスを計算した。
- Gray ら(2016)のインデックスに基づく創造性インデックスを適用し、著者およびジャンルごとの言語的独自性を評価した。
- Kintsch(2012)に基づく美的感受性インデックスを用いて、文学的作品における言語的芸術性を推定した。
- 意味的複雑性特徴を用いてテキスト分類および著者識別モデルを訓練し、正確性指標を用いて性能を評価した。
- すべての分析は、100人以上の著者と多様な文学形式を含むGutenberg Literary English Corpus (GLEC) で実施された。
実験結果
リサーチクエスチョン
- RQ1GLECにおいて、異なる文学ジャンル(例:演劇、詩、小説)は、トピック分布と感情プロファイルでどのように異なるか?
- RQ2文章内ばらつきと段階的距離は、テキストの文学的特徴を信頼できる指標としてどの程度有効か?
- RQ3Gray ら(2016)のインデックスに基づくと、どの文学ジャンルと著者が言語的創造性の面で最高水準に達しているか?
- RQ4意味的複雑性特徴は、テキスト分類および著者識別モデルの正確性を顕著に向上させることができるか?
- RQ5Kintsch(2012)の美的感受性インデックスに基づくと、ジェーン・オーステンの小説の中でどの作品が最も美的であると予測されるか?
主な発見
- 文章内ばらつきと段階的距離の指標に基づくと、演劇が最も文学的であるジャンルと特定された。続くのは詩と小説であった。
- 詩と演劇が創造性の面で最も高い水準に達しており、最も創造的な個々の著者たちは、ミルトン、ポープ、キーツ、バーリング、ワーズワースらの詩人であった。
- Kintsch(2012)の美的感受性インデックスを用いた予測では、小説『エマ』がジェーン・オーステンの作品の中で理論的に最も美的であるとされた。
- 意味的複雑性特徴は、テキスト分類および著者識別タスクにおいて75%から97%の予測正確性を達成した。
- 本研究では、文学的テキストにおける文学的特徴、創造性、美的感受性に関する検証済みの計算的ベンチマークを提供し、今後のデジタル・ヒューマニティーズおよび認知科学の研究の基盤を提供した。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。