[论文解读] Automatic quality evaluation and (semi-) automatic improvement of mixed models for OCR on historical documents.
本文提出了一种半自动方法,通过利用初始混合模型OCR系统生成的字符置信度和词汇性得分,指导新个性化OCR模型的训练,从而提升历史文献OCR的准确性。仅需40–80行人工校对的最小标注数据,该方法即可将字符错误率降至百分之几,显著减少人工转录工作量,同时无需事先进行字体特定建模即可克服排版多样性问题。
Good OCR results on historical documents rely on diplomatic transcriptions of printed material as ground truth which is both a scarce resource and time-consuming to generate. A strategy is proposed which starts from a mixed model trained on already available transcriptions from different centuries giving accuracies over 90% on a test set from the same period of time, overcoming the typography barrier of having to train individual models separately for each historical typeface. It is shown that both mean character confidence (as output by the OCR engine OCRopus) and lexicality (a measure of correctness of OCR tokens compared to a lexicon of modern wordforms taking historical spelling patterns into account, which can be calculated for any OCR engine) correlate with true accuracy determined from a comparison of OCR results with ground truth. These measures are then used to guide the training of new individual OCR models either using OCR prediction as pseudo ground truth (fully automatic method) or choosing a minimum set of hand-corrected lines as training material (manual method). Already 40-80 hand- corrected lines lead to OCR results with character error rates of only a few percent. This procedure minimizes the amount of ground truth production and does not depend on the previous construction of a specific typographic model.
研究动机与目标
- 解决用于训练历史文献准确OCR系统的外交转录数据稀缺且成本高昂的问题。
- 通过混合模型方法克服需为每种历史字体分别建模的限制。
- 开发一种最小化人工标注数据生产量的方法,同时在多样化历史书写系统上实现高OCR准确性。
- 评估OCR引擎输出(置信度和词汇性)是否能在无完整标注数据的情况下可靠预测真实OCR准确性。
提出的方法
- 在多个世纪的转录文本上训练初始混合模型OCR系统,在同期测试集上准确率超过90%。
- 在模型优化过程中,使用OCRopus生成的平均字符置信度作为OCR可靠性的代理指标。
- 通过将OCR生成的词元与现代词形词典进行比较,计算词汇性,同时考虑历史拼写变化。
- 利用置信度和词汇性得分,从OCR预测中选择高质量结果作为伪标注数据,用于训练新的独立模型。
- 应用一种人工方法,仅选择40–80行人工校对的文本对模型进行微调,从而降低对大规模标注数据的依赖。
- 通过置信度和词汇性指标迭代改进OCR模型,指导全自动和半自动训练流程。
实验结果
研究问题
- RQ1在无完整标注数据的情况下,OCR引擎输出的字符置信度和词汇性得分是否能可靠预测真实OCR准确性?
- RQ2从OCR输出中衍生的伪标注数据在多大程度上能提升历史文献上独立模型的性能?
- RQ3实现历史文本OCR低字符错误率,最少需要多少人工校对行?
- RQ4单一混合模型OCR系统是否能在不同历史时期和字体间良好泛化?
- RQ5置信度与词汇性度量的结合是否能减少OCR开发中对大规模人工转录的依赖?
主要发现
- 平均字符置信度和词汇性得分与真实OCR准确性具有强相关性,使得无需完整标注数据即可实现可靠的性能评估。
- 将OCR预测结果用作伪标注数据,可在极少人工干预下显著提升模型性能。
- 仅需40–80行人工校对的文本即可训练出字符错误率仅为百分之几的新OCR模型。
- 该方法有效克服了排版障碍,使在无需事先进行字体特定建模的情况下,对多样化历史字体实现高精度成为可能。
- 该方法减少了对大规模外交转录数据的依赖,使历史文献OCR开发更具可扩展性和效率。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。