Skip to main content
QUICK REVIEW

[论文解读] Anusaaraka: Overcoming the Language Barrier in India

Akshar Bharati, Vineet Chaitanya|ArXiv.org|Aug 7, 2003
Multilingual Education and PolicySocial Sciences被引用 20
一句话总结

Anusaaraka 提出了一种人机协同的机器翻译系统,通过将源文本转换为目标语言的音似和形似形式,实现在语义相近的印度语言之间的跨语言文本访问。该系统利用语言相近性,减轻了人类的翻译负担,使用户在约两周内即可掌握输出语言,再通过后期编辑实现语法准确性和风格调整。

ABSTRACT

The anusaaraka system makes text in one Indian language accessible in another Indian language. In the anusaaraka approach, the load is so divided between man and computer that the language load is taken by the machine, and the interpretation of the text is left to the man. The machine presents an image of the source text in a language close to the target language.In the image, some constructions of the source language (which do not have equivalents) spill over to the output. Some special notation is also devised. The user after some training learns to read and understand the output. Because the Indian languages are close, the learning time of the output language is short, and is expected to be around 2 weeks. The output can also be post-edited by a trained user to make it grammatically correct in the target language. Style can also be changed, if necessary. Thus, in this scenario, it can function as a human assisted translation system. Currently, anusaarakas are being built from Telugu, Kannada, Marathi, Bengali and Punjabi to Hindi. They can be built for all Indian languages in the near future. Everybody must pitch in to build such systems connecting all Indian languages, using the free software model.

研究动机与目标

  • 通过在印度语言之间实现文本访问,解决多语言印度的语义障碍。
  • 通过将语言处理任务交由机器承担,减轻人类翻译者的认知和语言负担。
  • 创建一个可扩展的、基于免费软件的框架,用于构建印度语言之间的翻译系统。
  • 利用印度语言之间的语言相似性,最小化用户理解机器生成输出所需的学习时间。
  • 通过后期编辑实现语法正确性和风格优化,使系统可作为人机协同翻译系统使用。

提出的方法

  • 系统在目标语言中生成源文本的视觉和语音近似形式,同时保留形态和句法结构。
  • 使用一种记号系统来表示源语言中在目标语言中缺乏直接对应形式的构造。
  • 输出设计为在视觉和语音上与目标语言高度接近,使用户在极短训练后即可快速理解。
  • 用户需接受训练,将输出理解为语音和拼写上的近似,而非字面翻译。
  • 经训练的用户对输出进行后期编辑,以纠正语法错误并提升风格,最终生成自然的目标语言文本。
  • 该方法采用免费软件模型实现,支持社区驱动的开发,并可扩展至所有印度语言。

实验结果

研究问题

  • RQ1机器生成的目标印度语言文本近似形式是否能帮助不熟悉该语言的用户实现快速理解?
  • RQ2印度语言之间的语言相似性在多大程度上可减少用户理解机器生成输出所需的学习时间?
  • RQ3该系统在通过后期编辑实现语法正确性和风格优化方面,对人机协同翻译的支持效果如何?
  • RQ4视觉和语音相似性在降低跨语言文本访问认知负担方面发挥何种作用?
  • RQ5免费软件模型能否持续支持印度多样化语言环境中多语言翻译系统的开发与扩展?

主要发现

  • 由于源语言与目标语言在语音和拼写上的相似性,用户可在约两周内学会阅读和理解机器生成的输出。
  • 该系统通过由机器承担大部分语言处理任务,成功减轻了人类翻译者的负担,使用户专注于解读与优化。
  • 经训练用户进行后期编辑后,输出在语法上正确且风格上自然得体。
  • 该方法具有可扩展性和可扩展性,Anusaaraka 系统目前已为泰卢固语、卡纳达语、马拉地语、孟加拉语和旁遮普语至印地语的转换开发完成。
  • 该系统作为人机协同翻译流水线运行良好,结合机器辅助与人工监督,显著提升了输出质量。
  • 免费软件模型支持协作开发,有望在未来短时间内为所有印度语言构建 Anusaaraka 系统。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。