Skip to main content
QUICK REVIEW

[논문 리뷰] A human-inspired recognition system for premodern Japanese historical documents

Anh Duc Le, Tarin Clanuwat|arXiv (Cornell University)|2019. 05. 14.
Handwritten Text Recognition Techniques참고 문헌 14인용 수 4
한 줄 요약

이 논문은 인간의 독서 방식을 모방하여 선형 시작을 감지하고, 문자를 순차적으로 스캔하며, 서체가 뒤섞이고 연결된 글자를 처리하는 인공지능 기반의 어텐션 기반 인코더-디코더 시스템을 제안한다. 이 시스템은 PRMU Kuzushiji 경쟁 데이터셋의 레벨 2에서 9.87%의 시퀀스 오류율을 기록했고, 레벨 3에서는 53.81%를 기록하여 참가한 모든 시스템보다 뛰어난 성능을 보였다.

ABSTRACT

Recognition of historical documents is a challenging problem due to the noised, damaged characters and background. However, in Japanese historical documents, not only contains the mentioned problems, pre-modern Japanese characters were written in cursive and are connected. Therefore, character segmentation based methods do not work well. This leads to the idea of creating a new recognition system. In this paper, we propose a human-inspired document reading system to recognize multiple lines of premodern Japanese historical documents. During the reading, people employ eyes movement to determine the start of a text line. Then, they move the eyes from the current character/word to the next character/word. They can also determine the end of a line or skip a figure to move to the next line. The eyes movement integrates with visual processing to operate the reading process in the brain. We employ attention-based encoder-decoder to implement this recognition system. First, the recognition system detects where to start a text line. Second, the system scans and recognize character by character until the text line is completed. Then, the system continues to detect the start of the next text line. This process is repeated until reading the whole document. We tested our human-inspired recognition system on the pre-modern Japanese historical document provide by the PRMU Kuzushiji competition. The results of the experiments demonstrate the superiority and effectiveness of our proposed system by achieving Sequence Error Rate of 9.87% and 53.81% on level 2 and level 3 of the dataset, respectively. These results outperform to any other systems participated in the PRMU Kuzushiji competition.

연구 동기 및 목표

  • 기존의 분할 기반 방법이 실패하는 손상된 역사적 문서에서 서체가 뒤섞이고 연결된 고대 일본어 문자를 인식하는 데 도전하는 것.
  • 눈동자의 움직임과 순차적 스캔 방식을 모방하여 복잡한 글자 스타일의 인식 성능을 향상시키기 위해 딥러닝 시스템에 인간의 독서 행동을 모델링하는 것.
  • 사전 분할 없이 선 시작을 감지하고, 한 줄씩 텍스트를 처리하며, 길이가 변하는 줄을 다룰 수 있는 시퀀스-투-시퀀스 인식 시스템을 개발하는 것.
  • 인간 유사 시각적 어텐션과 순차적 처리를 활용하여 PRMU Kuzushiji 경쟁에서 기존 시스템을 초월하는 성능을 달성하는 것.

제안 방법

  • 시스템은 문서 이미지를 시퀀스-투-시퀀스 방식으로 처리하기 위해 어텐션 기반 인코더-디코더 아키텍처를 사용한다.
  • 먼저 공간적 어텐션 메커니즘을 사용하여 각 텍스트 줄의 시작 위치를 감지한다.
  • 모델은 인간의 눈동자 움직임을 모방하여 문자 단위로 순차적으로 스캔하면서 인식을 수행한다.
  • 어텐션 메커니즘은 각 단계에서 관련된 이미지 영역에 집중함으로써 맥락 인식 능력을 향상시킨다.
  • 시스템은 한 줄씩 처리하며, 줄 끝을 감지하고 전체 문서를 읽을 때까지 다음 줄로 이동한다.
  • 아키텍처는 Kuzushiji 데이터셋에서 끝에서 끝까지 훈련되며, 오류율을 최소화하기 위해 시퀀스 수준의 손실 함수를 사용한다.

실험 결과

연구 질문

  • RQ1인간의 시각적 어텐션과 눈동자 움직임 패tern을 모방한 인식 시스템이 서체가 뒤섞이고 연결된 고대 일본어 글자에서 성능 향상에 기여할 수 있는가?
  • RQ2기존의 종단간 또는 분할 기반 접근 방식과 비교해 복잡한 역사적 일본어 문서에서 순차적이고 줄 단위로 처리하는 방식은 어떤가?
  • RQ3노이즈, 손상, 글자 스타일의 다양성 등 다양한 요소가 존재할 때 어텐션 메커니즘이 얼마나 정확도 향상에 기여하는가?
  • RQ4인간의 독서 행동을 모델링하면 도전적인 역사적 문서 벤치마크에서 더 나은 일반화 능력과 낮은 오류율을 달성할 수 있는가?

주요 결과

  • 제안된 시스템은 Kuzushiji 경쟁 데이터셋의 레벨 2에서 9.87%의 시퀀스 오류율(SER)을 기록하여 모든 다른 시스템을 압도했다.
  • 레벨 3에서는 SER가 53.81%로, 경쟁에 참가한 모든 참가자 중에서 가장 우수한 성능을 기록했다.
  • 인간 유사 어텐션 메커니즘 덕분에 문자 분할에 의존하지 않고도 서체가 뒤섞이고 연결된 문자를 견고하게 인식할 수 있었다.
  • 시스템의 성능는 인간 유사 순차적 스캔과 선 감지의 효과성을 입증했다.
  • 결과는 어텐션 기반 시퀀스 모델링이 높은 시각적 변동성이 있는 복잡한 역사적 글자 스타일에 매우 효과적이라는 것을 확인했다.
  • 시스템의 아키텍처는 고대 일본 수필의 다양한 문서 품질과 글자 스타일에 잘 일반화된다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.