[論文レビュー] Japanese/English Cross-Language Information Retrieval: Exploration of Query Translation and Transliteration
本稿では、技術文書向けに最適化された日本語/英語双方向クロスラングエージュアル情報検索(CLIR)システムを提案する。このシステムは、合成語翻訳、カタカナ語の発音表記化、頻度ベースの曖昧性解消を用いた省略語展開を統合しており、技術用語の翻訳精度を向上させることで、単言語IRと同等の性能を達成している。NACSISテストコレクションを用いた評価で、その有効性が検証された。
Cross-language information retrieval (CLIR), where queries and documents are in different languages, has of late become one of the major topics within the information retrieval community. This paper proposes a Japanese/English CLIR system, where we combine a query translation and retrieval modules. We currently target the retrieval of technical documents, and therefore the performance of our system is highly dependent on the quality of the translation of technical terms. However, the technical term translation is still problematic in that technical terms are often compound words, and thus new terms are progressively created by combining existing base words. In addition, Japanese often represents loanwords based on its special phonogram. Consequently, existing dictionaries find it difficult to achieve sufficient coverage. To counter the first problem, we produce a Japanese/English dictionary for base words, and translate compound words on a word-by-word basis. We also use a probabilistic method to resolve translation ambiguity. For the second problem, we use a transliteration method, which corresponds words unlisted in the base word dictionary to their phonetic equivalents in the target language. We evaluate our system using a test collection for CLIR, and show that both the compound word translation and transliteration methods improve the system performance.
研究の動機と目的
- 日本語と英語の間で技術文書を検索するクロスラングエージュアル情報検索(CLIR)の課題に対処すること。
- 標準的な二国語辞書では網羅されない技術用語のクエリ翻訳精度を向上させること。
- 合成技術用語の翻訳、カタカナ由来の借用語の発音表記化、省略語の展開を自動で行う手法を開発・統合すること。
- 日本語クエリと多言語技術要約からなる標準化されたNACSISテストコレクションを用いて、システムの性能を評価すること。
- これらの翻訳手法を組み合わせることで、技術分野におけるCLIRの有効性が顕著に向上することを示すこと。
提案手法
- 文書コレクションからの頻度統計を用いて、対象言語における基本語の組み合わせ確率に基づき翻訳を選択する合成語翻訳手法を提案した。
- 文字単位の対応により、日本語のカタカナ語を英語の発音的同等語にマッピングする発音表記手法を開発した。
- EDR技術用語辞書から基本語の翻訳を抽出することで、合成語用の二国語辞書を作成した。
- 英語コーパスから抽出した全形と略語を含む省略語辞書をシステムに統合した。
- 語の頻度統計を用いて翻訳および発音表記の曖昧性を解消し、自動的な曖昧性解消を可能にした。
- クエリ翻訳モジュールを単言語検索モデル(TF、対数TF、SMART)と統合し、標準的なIR手法を用いた性能評価を実施した。
実験結果
リサーチクエスチョン
- RQ1基本語の組み合わせ確率に基づく合成語翻訳は、技術文書向けCLIRの性能向上に寄与するか?
- RQ2日本語のカタカナ語の発音表記化は、日本語/英語CLIRにおける検索精度をどの程度向上させるか?
- RQ3省略語の展開(例:'IR' → 'information retrieval')を組み込むことで、システム性能にどのような影響を与えるか?
- RQ4合成語翻訳、発音表記化、省略語展開の複数の翻訳手法を組み合わせることで、CLIRの性能向上に相乗効果が得られるか?
- RQ5提案されたシステムは、技術文書検索において、単言語IRと同等の性能に到達できるか?
主な発見
- 合成語翻訳手法はベースラインを上回る検索性能を示し、他の手法と組み合わせることで特に顕著な向上が見られた。
- 発音表記手法は、カタカナ表記の技術用語において、性能向上に寄与した。
- 省略語の展開を組み込むことで、特定のケース(例:'LFG' → 'lexical functional grammar')で顕著なF1スコアの向上が観察された。
- 合成語翻訳、発音表記化、省略語展開の3つの手法を併用した場合、平均精度が日本語-日本語単言語IRと同等の水準に達した。
- 検索モジュールの強化(例:SMARTモデルの使用)は独立して性能向上をもたらしたため、より優れたクエリ翻訳の恩恵を、さらに強化する可能性があることが示された。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。