Skip to main content
QUICK REVIEW

[论文解读] Robust Quantification of Gender Disparity in Pre-Modern English Literature using Natural Language Processing

Akarsh Nagaraj, Mayank Kejriwal|arXiv (Cornell University)|Apr 12, 2022
Computational and Text Analysis Methods被引用 5
一句话总结

本研究利用 Project Gutenberg 语料库中的开源 NLP 工具,对 1850 至 1950 年间前现代英语文学作品进行稳健量化,分析性别差异,发现女性角色的出现频率显著低于男性角色。当控制作者性别后,这种差异有所减小,表明女性作者创作的作品中性别角色分布更为均衡。

ABSTRACT

Research has continued to shed light on the extent and significance of gender disparity in social, cultural and economic spheres. More recently, computational tools from the Natural Language Processing (NLP) literature have been proposed for measuring such disparity using relatively extensive datasets and empirically rigorous methodologies. In this paper, we contribute to this line of research by studying gender disparity, at scale, in copyright-expired literary texts published in the pre-modern period (defined in this work as the period ranging from the mid-nineteenth through the mid-twentieth century). One of the challenges in using such tools is to ensure quality control, and by extension, trustworthy statistical analysis. Another challenge is in using materials and methods that are publicly available and have been established for some time, both to ensure that they can be used and vetted in the future, and also, to add confidence to the methodology itself. We present our solution to addressing these challenges, and using multiple measures, demonstrate the significant discrepancy between the prevalence of female characters and male characters in pre-modern literature. The evidence suggests that the discrepancy declines when the author is female. The discrepancy seems to be relatively stable as we plot data over the decades in this century-long period. Finally, we aim to carefully describe both the limitations and ethical caveats associated with this study, and others like it.

研究动机与目标

  • 使用计算方法对前现代英语文学作品中的角色性别代表性进行定量测量。
  • 通过采用多种独立的角色出现频率测量方法和严谨的统计分析,确保方法论的稳健性。
  • 探究作者性别是否影响观察到的角色性别代表性差异。
  • 评估 1850 年至 1950 年这 100 年间性别角色代表性的随时间变化趋势。
  • 批判性地审视并公开披露方法论局限性和伦理问题,特别是关于非二元性别和性别非 conforming 个体的问题。

提出的方法

  • 使用开源的 Project Gutenberg 英语语料库,包含 1850 年至 1950 年间版权过期的文学作品。
  • 采用多阶段 NLP 流程,利用工业级开源 NLP 工具包,对角色和代词进行性别分类。
  • 应用三种不同的性别特定角色出现频率测量方法:(1) 命名角色的数量,(2) 角色提及次数,(3) 代词使用情况。
  • 通过在代表性样本中进行人工标注,估算性别分类的准确性,对姓名的分类结果达到近乎完美。
  • 在不同时间段和作者性别之间进行统计分析,以检验趋势和条件效应。
  • 公开发布所有代码和数据,以确保研究的可复现性,并接受研究社区的持续审查。

实验结果

研究问题

  • RQ1在前现代英语文学中,女性角色的出现频率是否显著低于男性角色?
  • RQ2当控制作者性别后,角色性别代表性差异是否减小?
  • RQ31850 年至 1950 年间,女性与男性角色出现频率的比值是否发生了显著变化?
  • RQ4方法选择和假设(如基于姓名的性别推断)在多大程度上影响了研究结果的可靠性?
  • RQ5在文学语料库中测量性别差异时,存在哪些伦理和方法论局限性,特别是针对非二元性别和性别非 conforming 个体?

主要发现

  • 在本研究使用的三种出现频率测量方法中,女性角色的出现频率显著低于男性角色。
  • 在控制作者性别后,角色性别代表性差异显著减小,但仍具有统计显著性。
  • 1850 年至 1950 年这 100 年间,男性与女性角色出现频率的比值保持相对稳定,女性角色代表性未随时间显著提升。
  • 本研究发现,当前 NLP 文献中尚无可靠方法可准确识别非二元或性别非 conforming 角色,凸显现有工具的关键空白。
  • 性别分类的准确率基于小规模人工标注样本估算,作者警告不应盲目依赖这些数据,而应在进一步验证前保持谨慎。
  • 作者强调,方法选择(包括假设设定和测量指标的选择)可能影响研究结果,主张未来应通过复制研究和提出替代方案来确保结果的稳健性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。