Skip to main content
QUICK REVIEW

[論文レビュー] Aksharantar: Open Indic-language Transliteration datasets and models for the Next Billion Users

Yash Madhani, Sushane Parthan|arXiv (Cornell University)|May 6, 2022
Natural Language Processing Techniques被引用数 5
ひとこと要約

本稿では、21の言語および12の script をカバーする2600万組のペアを含む、インド系言語の表記変換のための最大規模のオープンソースデータセット「Aksharantar」を紹介する。本稿では、Dakshinaベンチマークで15%の精度向上を達成し、新しい10万3千件のテストセットで強力なベースラインを確立する多言語モデル「IndicXlit」を提示しており、語の種類に応じた微細な分析を可能にしている。

ABSTRACT

Transliteration is very important in the Indian language context due to the usage of multiple scripts and the widespread use of romanized inputs. However, few training and evaluation sets are publicly available. We introduce Aksharantar, the largest publicly available transliteration dataset for Indian languages created by mining from monolingual and parallel corpora, as well as collecting data from human annotators. The dataset contains 26 million transliteration pairs for 21 Indic languages from 3 language families using 12 scripts. Aksharantar is 21 times larger than existing datasets and is the first publicly available dataset for 7 languages and 1 language family. We also introduce the Aksharantar testset comprising 103k word pairs spanning 19 languages that enables a fine-grained analysis of transliteration models on native origin words, foreign words, frequent words, and rare words. Using the training set, we trained IndicXlit, a multilingual transliteration model that improves accuracy by 15% on the Dakshina test set, and establishes strong baselines on the Aksharantar testset introduced in this work. The models, mining scripts, transliteration guidelines, and datasets are available at https://github.com/AI4Bharat/IndicXlit under open-source licenses. We hope the availability of these large-scale, open resources will spur innovation for Indic language transliteration and downstream applications. We hope the availability of these large-scale, open resources will spur innovation for Indic language transliteration and downstream applications.

研究の動機と目的

  • インド系言語向けの大規模で公開可能な表記変換データセットの不足に対処すること。
  • 複数のインド系言語系統と script をカバーする包括的でオープンソースのデータセットを構築すること。
  • ベンチマークおよび新しいテストセットで既存のベースラインを上回る多言語表記変換モデルを開発すること。
  • ネイティブ語、外国語、頻出語、希少語など、多様な語の種類における表記変換モデルの微細な評価を可能にすること。
  • 研究およびアプリケーション開発の促進を目的に、データセット、モデル、トレーニングスクリプト、ガイドラインをオープンアクセスで提供すること。

提案手法

  • データセットは、単語語彙および並列語彙を収集した後、ルールベースおよびニューラルフィルタリングを用いて有効な表記変換ペアを抽出することで構築された。
  • 低リソース言語および言語ペアの高品質なデータ収集のため、人間のアノテーターが活用され、言語的多様性と正確性が保証された。
  • 全 Aksharantar トレーニングセットを用いて、すべての言語で共通のサブワードボキャブラリーを採用した多言語シーケンス・ツー・シーケンスモデル「IndicXlit」が訓練された。
  • モデルは、インド系言語間の共通パターンを活用するため、クロスリンガルアテンションを備えたトランスフォーマーに基づくアーキテクチャを採用している。
  • 語の頻度、語源、 script 種別にわたる詳細な評価を可能にするために、10万3千の語ペアを含む新しいテストセット「Aksharantar testset」が作成された。
  • トレーニングおよび推論パイプラインは、完全なドキュメンテーションおよびオープンソースのライセンスのもとで公開された。

実験結果

リサーチクエスチョン

  • RQ1低リソースのインド系言語向けに、大規模かつ高品質な表記変換データを体系的に収集する方法は何か?
  • RQ2多様なインド系言語で訓練された多言語モデルは、言語系統や script 間でどの程度一般化できるか?
  • RQ3微細なセグメンテーションを用いた評価において、表記変換モデルは希少語、外国語、ネイティブ語に対してどの程度の性能を示すか?
  • RQ4オープンアクセスのデータセットおよびモデルは、Dakshina などの既存ベンチマークでの性能を顕著に向上させることができるか?
  • RQ5統一された表記変換トレーニング環境に、低リソースの言語および script を統合することで、どのような影響が生じるか?

主な発見

  • Aksharantar は、3つの言語系統および12の script をカバーする21のインド系言語で2600万組の表記変換ペアを含み、既存のデータセットの21倍の規模である。
  • このデータセットは、7つの言語および1つの言語系統について、初めて公開可能なリソースとして提供され、低リソース言語のカバレッジを顕著に拡大した。
  • Aksharantar で訓練された「IndicXlit」は、Dakshina テストセットにおいて、以前の最先端手法と比較して15%の絶対的精度向上を達成した。
  • Aksharantar テストセットにより、ネイティブ語、外国語、頻出語、希少語といった語の種別に応じた性能差の分析が可能になった。
  • 許容的なライセンスのもとでモデル、データセット、トレーニングスクリプトをオープンリリースしたことで、インド系言語 NLP 分野におけるイノベーションの加速が期待される。
  • 本研究では、多様なインド系言語で事前学習を行うことで、強固で一般化可能な表記変換性能が得られることを示した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。