Skip to main content
QUICK REVIEW

[论文解读] Synthetic data generation for Indic handwritten text recognition

Partha Pratim Roy, Akash Mohta|arXiv (Cornell University)|Apr 17, 2018
Handwritten Text Recognition Techniques参考文献 26被引用 6
一句话总结

本文提出了一种针对婆罗米系手写文本识别的合成数据生成方法,通过对手写数字文本施加受控的失真,以模拟自然手写变化。该方法显著提升了梵文和孟加拉文脚本在字符、数字和单词级别识别任务中的准确率,同时在拉丁文上也表现出色,证明了其在低资源手写识别系统中的广泛适用性。

ABSTRACT

This paper presents a novel approach to generate synthetic dataset for handwritten word recognition systems. It is difficult to recognize handwritten scripts for which sufficient training data is not readily available or it may be expensive to collect such data. Hence, it becomes hard to train recognition systems owing to lack of proper dataset. To overcome such problems, synthetic data could be used to create or expand the existing training dataset to improve recognition performance. Any available digital data from online newspaper and such sources can be used to generate synthetic data. In this paper, we propose to add distortion/deformation to digital data in such a way that the underlying pattern is preserved, so that the image so produced bears a close similarity to actual handwritten samples. The images thus produced can be used independently to train the system or be combined with natural handwritten data to augment the original dataset and improve the recognition system. We experimented using synthetic data to improve the recognition accuracy of isolated characters and words. The framework is tested on 2 Indic scripts - Devanagari (Hindi) and Bengali (Bangla), for numeral, character and word recognition. We have obtained encouraging results from the experiment. Finally, the experiment with Latin text verifies the utility of the approach.

研究动机与目标

  • 解决梵文和孟加拉文等低资源婆罗米系手写文本标注数据集稀缺的问题。
  • 通过生成模拟自然手写的合成样本,降低真实手写数据收集的成本与工作量。
  • 通过数据增强提升孤立字符与单词识别系统的识别性能。
  • 验证该方法在多种脚本(包括拉丁文)上的泛化能力。
  • 在引入逼真失真时,保持原始文本模式不变,确保合成数据对训练仍具实用性。

提出的方法

  • 通过在在线来源获取的干净数字文本图像上施加受控的几何与光度失真,生成合成数据。
  • 失真包括扭曲、旋转、缩放和噪声添加,以模拟自然手写的变化。
  • 该方法确保在视觉变化下,原始文本内容与结构在语义上保持完整。
  • 生成的合成图像可独立使用,或与真实手写数据结合用于训练识别模型。
  • 该框架在梵文和孟加拉文脚本的孤立字符与单词识别任务上进行了评估。
  • 消融研究与跨脚本评估(拉丁文)验证了合成数据的鲁棒性与可迁移性。

实验结果

研究问题

  • RQ1通过数字文本失真生成的合成数据,能否提升低资源婆罗米系手写文本的识别准确率?
  • RQ2合成数据增强在孤立字符与单词识别任务上的性能提升程度如何?
  • RQ3该方法在不同脚本(包括非婆罗米系脚本如拉丁文)上的泛化能力如何?
  • RQ4合成数据在引入逼真手写风格变化的同时,是否保持了语义内容的完整性?
  • RQ5当与真实手写数据结合时,合成数据在训练识别模型中的相对贡献如何?

主要发现

  • 所提出的合成数据生成方法在字符、数字和单词识别任务上,显著提升了梵文和孟加拉文脚本的识别准确率。
  • 将合成数据与真实手写数据结合使用,性能优于仅使用真实数据,证明了数据增强的有效性。
  • 该方法在拉丁文脚本上也表现出良好泛化能力,证实其在婆罗米系脚本之外的适用性。
  • 合成样本在视觉上与真实手写样本高度相似,适合用于训练识别系统。
  • 该方法有效缓解了低资源手写识别场景下的数据稀缺问题。
  • 结果表明,基于失真的合成数据生成是一种可行且高效的替代昂贵数据收集的方法。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。