Skip to main content
QUICK REVIEW

[论文解读] Corpus of Chinese Dynastic Histories: Gender Analysis over Two Millennia

Sergey A. Zinin, Yang Xu|arXiv (Cornell University)|May 18, 2020
Gender Studies in Language被引用 5
一句话总结

本文介绍了一个开源语料库,包含24部中国历代正史,涵盖2000年来的文言文文献,旨在实现对文言文中性别语言的计算分析。通过关键词分析性别特定词汇,研究揭示了男性称谓长期占主导地位的显著历史稳定性,并生成了对历时语义学研究具有实用价值的语义表征。

ABSTRACT

Chinese dynastic histories form a large continuous linguistic space of approximately 2000 years, from the 3rd century BCE to the 18th century CE. The histories are documented in Classical (Literary) Chinese in a corpus of over 20 million characters, suitable for the computational analysis of historical lexicon and semantic change. However, there is no freely available open-source corpus of these histories, making Classical Chinese low-resource. This project introduces a new open-source corpus of twenty-four dynastic histories covered by Creative Commons license. An original list of Classical Chinese gender-specific terms was developed as a case study for analyzing the historical linguistic use of male and female terms. The study demonstrates considerable stability in the usage of these terms, with dominance of male terms. Exploration of word meanings uses keyword analysis of focus corpora created for genderspecific terms. This method yields meaningful semantic representations that can be used for future studies of diachronic semantics.

研究动机与目标

  • 为解决当前计算语言学中缺乏免费获取、开源的文言文历代正史语料库的问题,这些语料目前属于低资源类型。
  • 制定一份标准化的文言文性别特定词汇列表,以实现系统性的语言学分析。
  • 探究过去两千年历史文献中男性与女性称谓使用情况的历时性模式。
  • 通过聚焦语料的关键词分析,生成性别称谓的语义表征。
  • 通过可重复使用、已授权的语料库,为未来关于语义演变与文言文历史词汇的研究提供支持。

提出的方法

  • 从24部中国历代正史中汇编出超过2000万字符的语料库,涵盖公元前3世纪至公元18世纪的时期。
  • 应用新创建的文言文性别特定词汇列表,识别并提取语料库中相关的词汇项目。
  • 使用关键词语境(KWIC)和关键词分析技术,考察男性与女性称谓的语义特征。
  • 围绕性别特定词汇创建聚焦语料,以实现对比性语义分析。
  • 采用计算语言学方法,建模语义随时间的稳定性与变化。
  • 以知识共享许可协议发布语料库,确保学术研究的开放获取与可重用性。

实验结果

研究问题

  • RQ1过去两千年中,文言文历代正史中性别特定词汇的使用方式如何演变?
  • RQ2在历史文献中,男性与女性称谓在词汇与句法使用上的语义稳定性在多大程度上存在?
  • RQ3通过关键词分析考察性别称谓的语境使用时,会浮现哪些语义模式?
  • RQ4性别称谓的聚焦语料如何产生对历时语义学研究具有意义的表征?
  • RQ5文言文历史叙述中性别语言的结构与词汇特征是什么?

主要发现

  • 该语料库在所有朝代时期均显示出男性称谓对女性称谓的持续主导,表明长期存在的语言与文化不对称性。
  • 文言文中的性别特定词汇在使用模式上表现出显著的稳定性,2000年跨度内语义漂移极小。
  • 对聚焦语料的关键词分析成功生成了性别称谓的连贯语义表征,证实其在历时研究中的实用性。
  • 本研究证实,计算方法能够有效提取并分析文言文等低资源历史语言中的语义模式。
  • 语料库的开源化使得对历史中文文本中性别与语言的大规模、可复现分析成为可能。
  • 该项目为未来关于语义演变、性别表征及东亚文本传统中历史语篇的研究奠定了基础性资源。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。