[论文解读] Transcribing Medieval Manuscripts for Machine Learning
本文提出了一种使用手写文本识别(HTR)技术转录中世纪拉丁手稿的框架,以支持机器学习任务,强调针对学术研究需求定制的转录方案。该研究展示了如何通过HTR生成的转录文本实现对文本见证中抄写员差异、缩写及书写风格转换的可扩展分析。
This article focuses on the transcription of medieval manuscripts. Whereas problems of transcription have long interested medievalists, few workable options in the era of printed editions were available besides normalisation. The automation of this process, known as handwritten text recognition (HTR), has made new kinds of digital text creation possible, but also has foregrounded the necessity of theorising transcription in our scholarly practices. We reflect here on different notions of transcription against the backdrop of changing text technologies. Moreover, drawing on our own research on medieval Latin Bibles, we present general guidelines for customizing transcription schemes, arguing that they must be designed with specific research questions and scholarly end use in mind. Since we are particularly interested in the scribal contribution to the production of codices, our transcription guidelines aim to capture abbreviations and orthographic variation between different textual witnesses for downstream machine learning tasks. In the final section of the article, we discuss a few examples of how the HTR-created transcriptions allow us to address new questions at scale in medieval manuscripts, such as textual variance across witnesses, the prediction of a change in scribal hands within a single manuscript as well as the profiling of individual and regional scribal characteristics.
研究动机与目标
- 解决在机器学习时代,中世纪手稿系统化、理论指导的转录方法存在的空白。
- 开发可自定义的转录方案,以保留手稿见证中特有的古文字学与拼写变体,服务于下游HTR与机器学习应用。
- 通过捕捉不同手稿见证中缩写与拼写多样性的特征,支持对抄写员贡献的研究。
- 利用HTR生成的转录文本,实现对文本变异与个体抄写员特征的大规模分析。
- 为学者提供实用指南,帮助其设计与特定研究问题及数字学术目标相一致的转录方案。
提出的方法
- 设计保留特定手稿见证中拼写变体与缩写特征的转录方案。
- 将HTR模型应用于中世纪拉丁圣经的数字化图像,生成机器可读的转录文本。
- 定制转录处理流程,以反映学术优先事项,例如追踪单份手稿中抄写风格的变化。
- 利用HTR输出作为训练数据,训练专注于检测抄写员转换与区域抄写员特征的机器学习模型。
- 将古文字学特征(如连笔与缩写)整合进结构化转录格式,以支持计算分析。
- 通过与外交式转录及专家古文字学评估进行对比,验证转录的保真度。
实验结果
研究问题
- RQ1如何设计转录方案,以保留古文字学与拼写变体,从而支持机器学习应用?
- RQ2HTR生成的转录文本在多大程度上能够检测单份手稿中抄写员书写风格的变化?
- RQ3基于HTR的转录文本能否揭示多个手稿见证之间区域性和个体性抄写特征?
- RQ4在转录中包含缩写与拼写变体如何提升下游机器学习任务的准确性?
- RQ5转录选择对中世纪拉丁圣经中文本变异大规模分析有何影响?
主要发现
- 保留缩写与拼写变体的定制化转录方案,显著提升了HTR输出在古文字学分析中的保真度。
- HTR生成的转录文本使大规模检测单份手稿中抄写员书写风格的转换成为可能,此前此类分析需依赖人工检查。
- 基于HTR转录文本训练的机器学习模型,能够成功依据拼写与缩写模式识别个体与区域性的抄写员特征。
- 该方法使对多个手稿见证中文本变异的大规模比较成为可能,揭示了此前未被发现的抄写倾向。
- 围绕特定研究问题设计的转录方案,相较于通用的归一化方法,在下游机器学习任务中产生更准确、更有意义的结果。
- 将HTR与学术转录实践相结合,开启了数字手稿研究的新形式,尤其在分析抄写员能动性与文本传承方面具有重要意义。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。