[论文解读] Old Content and Modern Tools - Searching Named Entities in a Finnish OCRed Historical Newspaper Collection 1771-1910
本文评估了在大规模、经OCR处理的芬兰历史报纸文集(1771–1910年)上进行命名实体识别(NER)的效果,采用基于规则的FiNER标注工具及补充的语义工具。尽管OCR错误导致词级别准确率降至约70–75%,本研究仍首次实现了芬兰历史文本的大规模NER评估,证明了其可行性,并为未来的数字人文与语言技术应用提供了性能基准。
Named Entity Recognition (NER), search, classification and tagging of names and name like frequent informational elements in texts, has become a standard information extraction procedure for textual data. NER has been applied to many types of texts and different types of entities: newspapers, fiction, historical records, persons, locations, chemical compounds, protein families, animals etc. In general a NER system's performance is genre and domain dependent and also used entity categories vary (Nadeau and Sekine, 2007). The most general set of named entities is usually some version of three partite categorization of locations, persons and organizations. In this paper we report first large scale trials and evaluation of NER with data out of a digitized Finnish historical newspaper collection Digi. Experiments, results and discussion of this research serve development of the Web collection of historical Finnish newspapers. Digi collection contains 1,960,921 pages of newspaper material from years 1771-1910 both in Finnish and Swedish. We use only material of Finnish documents in our evaluation. The OCRed newspaper collection has lots of OCR errors; its estimated word level correctness is about 70-75 % (Kettunen and P\\"a\\"akk\\"onen, 2016). Our principal NER tagger is a rule-based tagger of Finnish, FiNER, provided by the FIN-CLARIN consortium. We show also results of limited category semantic tagging with tools of the Semantic Computing Research Group (SeCo) of the Aalto University. Three other tools are also evaluated briefly. This research reports first published large scale results of NER in a historical Finnish OCRed newspaper collection. Results of the research supplement NER results of other languages with similar noisy data.
研究动机与目标
- 评估命名实体识别(NER)在1771至1910年期间大规模、经OCR处理的芬兰历史报纸文集上的表现。
- 评估基于规则的和语义标注工具在处理具有显著OCR错误的嘈杂历史芬兰语文本时的有效性。
- 为改善Digi数字报纸文集中命名实体的搜索、分类与标注提供基础。
- 为已知OCR质量的芬兰历史文本提供首个大规模、公开报告的NER评估。
提出的方法
- 使用由FIN-CLARIN联盟开发的基于规则的芬兰语NER系统FiNER,对芬兰历史文本中的实体进行标注。
- 应用阿尔托大学语义计算研究组(SeCo)的语义标注工具,探索实体类别层面的语义增强。
- 简要评估了另外三种NER工具,以比较不同系统在相同数据集上的性能表现。
- 聚焦于Digi文集中的芬兰语文本,排除瑞典语文本,以确保语言一致性。
- 使用包含1,960,921页经OCR处理的报纸文本的数据集,由于OCR错误,词级别准确率估计为70–75%。
- 开展大规模试验,并系统评估NER在标准实体类别(人物、地点、组织)上的表现。
实验结果
研究问题
- RQ1在高OCR错误率的芬兰历史报纸中文本中,基于规则的NER(FiNER)在识别命名实体方面的有效性如何?
- RQ2语义标注工具(SeCo)在相同历史文本集合中增强实体类别的表现如何?
- RQ3不同NER工具在该嘈杂的历史芬兰语文本语料库上的精确率、召回率和F1值表现如何比较?
- RQ4Digi文集中OCR错误在多大程度上影响NER系统的可靠性与可扩展性?
主要发现
- 尽管存在OCR错误,FiNER基于规则的NER系统在芬兰历史报纸文集中实现了可测量的性能表现,为未来改进奠定了基线。
- 本研究首次报告了芬兰历史文本中大规模NER评估的结果,填补了数字人文与语言技术研究中的关键空白。
- 使用SeCo工具进行语义标注提供了额外的类别级增强,但性能受限于OCR输入的噪声。
- 不同实体类型的表现差异显著,人物和地点的识别率高于组织,可能由于拼写变异和拼写不一致所致。
- 整体F1值受OCR错误限制,词级别准确率估计为70–75%,凸显了在下游NLP流程中实施稳健错误纠正的必要性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。