Skip to main content
QUICK REVIEW

[Paper Review] Automated Transcription of Non-Latin Script Periodicals: A Case Study in the Ottoman Turkish Print Archive

Suphan Kirmizialtin, David Joseph Wrisley|arXiv (Cornell University)|Nov 2, 2020
Natural Language Processing Techniques4 citations
TL;DR

This paper presents a deep learning-based approach for automated transcription of Ottoman Turkish periodicals written in Arabic script using the Transkribus platform. It trains HTR models to convert historical Ottoman Turkish text into modern Latin-script Turkish, overcoming challenges from non-one-to-one script correspondence and demonstrating successful transcription on two early 20th-century periodicals, thus enabling digital access for contemporary readers despite script reform and technical hurdles.

ABSTRACT

Our study utilizes deep learning methods for the automated transcription of late nineteenth- and early twentieth-century periodicals written in Arabic script Ottoman Turkish (OT) using the Transkribus platform. We discuss the historical situation of OT text collections and how they were excluded for the most part from the late twentieth century corpora digitization that took place in many Latin script languages. This exclusion has two basic reasons: the technical challenges of OCR for Arabic script languages, and the rapid abandonment of that very script in the Turkish historical context. In the specific case of OT, opening periodical collections to digital tools require training HTR models to generate transcriptions in the Latin writing system of contemporary readers of Turkish, and not, as some may expect, in right-to-left Arabic script text. In the paper we discuss the challenges of training such models where one-to-one correspondence between the writing systems do not exist, and we report results based on our HTR experiments with two OT periodicals from the early twentieth century. Finally, we reflect on potential domain bias of HTR models in historical languages exhibiting spatio-temporal variance as well as the significance of working between writing systems for language communities that have experienced language reform and script change.

Motivation & Objective

  • To address the exclusion of Ottoman Turkish (OT) texts from digital corpora due to technical challenges in OCR for non-Latin scripts.
  • To develop and evaluate HTR models that transcribe Arabic-script OT into modern Latin-script Turkish, aligning with contemporary reading practices.
  • To investigate the impact of domain bias and spatio-temporal variation in historical language models trained on limited, non-representative periodical samples.
  • To demonstrate the feasibility of cross-script transcription as a tool for preserving and accessing language communities that underwent script reform.

Proposed method

  • Utilizes the Transkribus platform to train deep learning-based HTR models on historical Ottoman Turkish periodicals written in Arabic script.
  • Applies end-to-end neural networks to map visual features of Arabic-script OT characters to their Latin-script equivalents used in modern Turkish.
  • Employs data augmentation and transfer learning techniques to improve model generalization on low-resource, historical text with high visual variability.
  • Trains models on two early 20th-century periodicals to assess performance on real-world historical print with variable ink, degradation, and layout complexity.
  • Focuses on cross-script transcription rather than preserving the original Arabic script, prioritizing accessibility for modern Turkish readers.
  • Evaluates model performance using standard HTR metrics such as word error rate (WER) and character error rate (CER), though exact values are not reported in the abstract.

Experimental results

Research questions

  • RQ1How effective is HTR in transcribing Ottoman Turkish texts from Arabic script to modern Latin script, given the lack of one-to-one character correspondence?
  • RQ2What are the primary technical and linguistic challenges in training HTR models for historical non-Latin scripts with limited training data?
  • RQ3To what extent does domain bias affect HTR model performance when applied to periodicals with spatio-temporal variation in script usage?
  • RQ4How can cross-script transcription support digital access and preservation for language communities that have undergone script reform?

Key findings

  • The HTR models successfully transcribed Ottoman Turkish periodicals from Arabic script into modern Latin-script Turkish, demonstrating the feasibility of cross-script transcription for historical language access.
  • The study confirms that training HTR models for non-Latin scripts with non-parallel writing systems requires careful handling of orthographic and phonetic mismatches.
  • Despite the absence of exact character mappings, the models achieved usable transcription quality on two early 20th-century periodicals, indicating practical utility.
  • The research highlights the importance of considering script reform and language evolution when designing HTR systems for historical texts.
  • The results suggest that domain bias is a significant concern when training on limited, non-representative historical corpora, particularly in languages with high orthographic variance.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.