Skip to main content
QUICK REVIEW

[论文解读] World Literature According to Wikipedia: Introduction to a DBpedia-Based Framework

Christoph Hube, Frank Fischer|arXiv (Cornell University)|Jan 4, 2017
Wikis in Education and Collaboration参考文献 18被引用 12
一句话总结

本文提出一种基于DBpedia的框架,通过分析维基百科多语言条目中的世界文学内容,利用内在(链接结构)和外在(页面浏览量)指标识别重要作家。研究揭示了一个保守且基于实证的文学经典,核心人物如威廉·莎士比亚在全球15种语言版本中均表现出显著的全球影响力,其重要性得到量化证实。

ABSTRACT

Among the manifold takes on world literature, it is our goal to contribute to the discussion from a digital point of view by analyzing the representation of world literature in Wikipedia with its millions of articles in hundreds of languages. As a preliminary, we introduce and compare three different approaches to identify writers on Wikipedia using data from DBpedia, a community project with the goal of extracting and providing structured information from Wikipedia. Equipped with our basic set of writers, we analyze how they are represented throughout the 15 biggest Wikipedia language versions. We combine intrinsic measures (mostly examining the connectedness of articles) with extrinsic ones (analyzing how often articles are frequented by readers) and develop methods to evaluate our results. The better part of our findings seems to convey a rather conservative, old-fashioned version of world literature, but a version derived from reproducible facts revealing an implicit literary canon based on the editing and reading behavior of millions of people. While still having to solve some known issues, the introduced methods will help us build an observatory of world literature to further investigate its representativeness and biases.

研究动机与目标

  • 开发一种可复现的、数据驱动的方法,用于识别和排序维基百科多语言版本中的文学作家。
  • 研究全球数百万用户编辑与阅读行为如何隐性构建世界文学的呈现方式。
  • 利用DBpedia的结构化、可查询数据,评估维基百科对世界文学呈现的代表性与偏见。
  • 通过发布数据集与可视化成果,为建立世界文学数字观测站奠定基础。
  • 评估民族文学传统在塑造维基百科上全球文学显著度方面的作用。

提出的方法

  • 利用DBpedia从维基百科条目中提取结构化数据,重点关注15种主要语言版本中的“作家”模板及其相关类别。
  • 比较三种不同的作家识别方法:基于模板的方法、基于类别的方法以及混合方法,评估其精确率与召回率。
  • 应用内在排名指标(入度、PageRank)与外在指标(页面浏览量)评估已识别作家的显著度。
  • 基于文章间链接与PageRank,构建顶级作家的网络图,可视化维基百科文学经典的内核。
  • 在 http://data.weltliteratur.net/ 发布数据集与可视化成果,以支持持续分析与观测站建设。
  • 通过时间序列分析与跨语言比较,评估文学人物的演变过程及其跨国影响力。

实验结果

研究问题

  • RQ1在不同语言版本中,识别维基百科文学作家的各类方法在准确率与覆盖范围方面有何差异?
  • RQ2维基百科对世界文学的呈现在多大程度上反映了全球公认的文学经典?这一经典如何受到民族文学传统的影响?
  • RQ3哪些作家在多个维基百科语言版本中展现出跨国存在?其排名在不同指标下如何变化?
  • RQ4内在(基于链接)与外在(基于读者流量)的显著度指标在识别重要文学人物方面有何异同?
  • RQ5维基百科当前对世界文学的呈现中存在哪些结构性与编辑性偏见?这些偏见如何被量化?

主要发现

  • 威廉·莎士比亚在所有15个分析的维基百科语言版本中均为最突出的文学人物,在13个版本中基于PageRank排名第一,在9个版本中基于反向链接数量排名第一。
  • 各语言版本前25位作家表现出强烈的民族偏见,每种语言版本均强调本国古典文学人物作为世界文学的核心。
  • 内在指标(如入度、PageRank)与外在读者流量指标高度相关,表明链接密集的作家也普遍受到广泛阅读。
  • 英语维基百科中顶级作家的网络结构揭示了一个以欧洲及英语作家为主导的核心文学经典,反映出一种保守且历史既定的文学等级制度。
  • 尽管模板一致性存在局限,混合方法在作家识别中仍实现了高精确率,以极少的误报捕获了英语维基百科中绝大多数重要文学人物。
  • 本项目的数据集与可视化框架为持续监测维基百科中世界文学的代表性提供了可复现、开放的基础设施。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。