Skip to main content
QUICK REVIEW

[论文解读] OCR ACCURACY IMPROVEMENT ON DOCUMENT IMAGES THROUGH A NOVEL PRE-PROCESSING APPROACH

A. El Harraj, Naoufal Raissouni|arXiv (Cornell University)|Sep 11, 2015
Handwritten Text Recognition Techniques被引用 12
一句话总结

本文提出了一种新颖的非参数化、无监督预处理流程,以提升由数码相机和移动设备拍摄的退化文档图像的OCR准确率。通过依次应用局部亮度/对比度调整、优化的灰度转换、非锐化掩模以及全局二值化,该方法显著提升了标准基准测试中的文本检测与OCR性能。

ABSTRACT

Digital camera and mobile document image acquisition are new trends arising in the world of Optical Character Recognition and text detection. In some cases, such process integrates many distortions and produces poorly scanned text or text-photo images and natural images, leading to an unreliable OCR digitization. In this paper, we present a novel nonparametric and unsupervised method to compensate for undesirable document image distortions aiming to optimally improve OCR accuracy. Our approach relies on a very efficient stack of document image enhancing techniques to recover deformation of the entire document image. First, we propose a local brightness and contrast adjustment method to effectively handle lighting variations and the irregular distribution of image illumination. Second, we use an optimized greyscale conversion algorithm to transform our document image to greyscale level. Third, we sharpen the useful information in the resulting greyscale image using Un-sharp Masking method. Finally, an optimal global binarization approach is used to prepare the final document image to OCR recognition. The proposed approach can significantly improve text detection rate and optical character recognition accuracy. To demonstrate the efficiency of our approach, an exhaustive experimentation on a standard dataset is presented

研究动机与目标

  • 解决由移动设备拍摄文档图像中光照变化和失真导致的OCR不准确问题。
  • 开发一种非参数化、无监督的预处理方法,以在无需标注数据或模型训练的情况下提升图像质量。
  • 通过系统性地增强图像对比度、锐度和二值化,提升文本检测与OCR准确率。
  • 在真实成像条件下,于标准数据集上验证该方法的有效性。

提出的方法

  • 应用局部亮度与对比度调整技术,以校正文档图像中不均匀的光照与光照变化。
  • 采用优化的灰度转换算法,将输入的彩色图像转换为高质量的灰度表示。
  • 对灰度图像应用非锐化掩模,以增强边缘细节并锐化文本特征。
  • 采用最优全局二值化方法,将增强后的灰度图像转换为适合OCR处理的二值格式。
  • 将整个预处理流程作为一系列图像增强操作的堆叠顺序执行。
  • 该方法完全无监督且非参数化,仅依赖图像统计信息,无需外部参数或训练数据。

实验结果

研究问题

  • RQ1如何提升由移动设备或数码相机拍摄、常受光照与失真影响的文档图像的OCR准确率?
  • RQ2哪些预处理技术的组合能对OCR就绪的文档图像实现最有效的增强?
  • RQ3在缺乏源图像先验知识的情况下,非参数化且无监督的方法能否在处理多样化图像退化时优于现有方法?

主要发现

  • 所提出的预处理流程显著提升了在光照不均与低对比度文档图像上的文本检测率。
  • 局部对比度调整、优化的灰度转换、非锐化掩模与全局二值化的组合,在OCR准确率上带来了可测量的提升。
  • 该方法在无需训练数据或参数调优的情况下,对多样化的真实世界文档图像表现出稳健性能。
  • 在标准数据集上的实验验证了该方法在复杂成像条件下提升OCR可靠性的有效性。
  • 该方法的无监督与非参数化特性确保了其在各类文档类型与采集设备中的广泛适用性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。