[論文レビュー] Development of a Multi-User Recognition Engine for Handwritten Bangla Basic Characters and Digits
この論文では、Tesseract OCRエンジンを用いて、手書きバンガラ文字および数字のマルチユーザー認識エンジンを提示する。各ユーザーのデータセット(1ユーザーあたり919〜928サンプル)で訓練されたシステムは、文字レベルで92.15%、数字レベルで97.37%の正確性を達成し、個々のユーザーの認識正確性は90.66%、91.66%、96.87%であった。
The objective of the paper is to recognize handwritten samples of basic Bangla characters using Tesseract open source Optical Character Recognition (OCR) engine under Apache License 2.0. Handwritten data samples containing isolated Bangla basic characters and digits were collected from different users. Tesseract is trained with user-specific data samples of document pages to generate separate user-models representing a unique language-set. Each such language-set recognizes isolated basic Bangla handwritten test samples collected from the designated users. On a three user model, the system is trained with 919, 928 and 648 isolated handwritten character and digit samples and the performance is tested on 1527, 14116 and 1279 character and digit samples, collected form the test datasets of the three users respectively. The user specific character/digit recognition accuracies were obtained as 90.66%, 91.66% and 96.87% respectively. The overall basic character-level and digit level accuracy of the system is observed as 92.15% and 97.37%. The system fails to segment 12.33% characters and 15.96% digits and also erroneously classifies 7.85% characters and 2.63% on the overall dataset.
研究の動機と目的
- 手書きバンガラ基本文字および数字のスケーラブルでユーザー固有の認識システムの開発。
- 異なるユーザー間での筆跡のばらつきという課題への対処。
- 最小限の変更でオープンソースのTesseract OCRエンジンをバンガラ文字用に適応。
- 多様なユーザーから収集した分離された手書きサンプルを用いたシステムの性能評価。
- 認識パイプラインにおけるセグメンテーションエラーおよび分類エラーの定量的評価。
提案手法
- 3人の異なるユーザーから、バンガラ基本文字および数字の分離された手書きサンプルを収集した。
- 各ユーザーの固有の筆跡データを用いて、個別のTesseract言語モデルを訓練した。
- ユーザー固有のトレーニングデータセット(919、928、648サンプル)を用いて、個々のユーザー用モデルを生成した。
- 各ユーザーごとに1,527、14,116、1,279サンプルの独立したテストセットを用いて認識性能をテストした。
- Tesseractの既存OCRパイプラインに言語セットカスタマイズを適用し、ユーザー固有の筆跡のばらつきに対応した。
- 全データセット上で認識正確性、セグメンテーション失敗率、誤分類率を測定した。
実験結果
リサーチクエスチョン
- RQ1Tesseract OCRエンジンを用いて、手書きバンガラ文字および数字のマルチユーザー認識エンジンを効果的に構築できるか?
- RQ2ユーザー固有のトレーニングは、汎用モデルと比較して認識正確性をどのように向上させるか?
- RQ3このシステムにおける文字セグメンテーションおよび誤分類の失敗率はどの程度か?
- RQ4筆跡のスタイルが異なるユーザー間で、認識正確性はどのように変動するか?
- RQ5文字および数字認識正確性という観点から、このシステムの全体的な性能はどの程度か?
主な発見
- 全ユーザーにわたる文字レベルの認識正確性は92.15%であった。
- 数字レベルの認識正確性は97.37%に達し、数字文字における優れた性能を示した。
- 個々のユーザーの認識正確性は90.66%、91.66%、96.87%であり、ユーザー依存の性能差を示した。
- システムは12.33%の文字と15.96%の数字を正しくセグメンテーションできず、筆跡のセグメンテーションに課題があることを示した。
- 文字の7.85%と数字の2.63%が誤分類され、全体の誤り率に寄与した。
- ユーザー固有モデルの使用により、単一の汎用モデルと比較して認識正確性が顕著に向上した。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。