Skip to main content
QUICK REVIEW

[论文解读] Incorporating Discourse Aspects in English -- Polish MT: Towards Robust Implementation

Małgorzata E. Styś, Stefan Zemke|ArXiv.org|Oct 15, 1995
Natural Language Processing Techniques被引用 14
一句话总结

本文提出了一种用于英波机器翻译的语篇意识机器翻译系统,该系统采用扩展的中心理论来建模信息结构,并预测波兰语中的成分顺序。通过整合语用、句法和统计线索(包括定指性、代词指代和显著性分级),该系统根据语篇突出性重新排列波兰语成分,从而生成更具交际恰当性的翻译,显著提升了基线系统在流畅性和连贯性方面的表现。

ABSTRACT

The main aim of translation is an accurate transfer of meaning so that the result is not only grammatically and lexically correct but also communicatively adequate. This paper stresses the need for discourse analysis the aim of which is to preserve the communicative meaning in English--Polish machine translation. Unlike English, which is a positional language with word order grammatically determined, Polish displays a strong tendency to order constituents according to their degree of salience, so that the most informationally salient elements are placed towards the end of the clause regardless of their grammatical function. The Centering Theory developed for tracking down given information units in English and the Theory of Functional Sentence Perspective predicting informativeness of subsequent constituents provide theoretical background for this work. The notion of {\em center} is extended to accommodate not only for pronominalisation and exact reiteration but also for definiteness and other center pointing constructs. Center information is additionally graded and applicable to all primary constituents in a given utterance. This information is used to order the post-transfer constituents correctly, relying on statistical regularities and some syntactic clues.

研究动机与目标

  • 解决现有机器翻译系统在句法层面准确之外无法保持交际意义的缺陷。
  • 使用带有分级中心值的扩展中心理论,在英语中建模语篇级信息结构。
  • 开发一种基于规则的稳健系统,根据语篇突出性和显著性对波兰语成分进行排序。
  • 将句法、语义和统计线索整合到统一框架中,用于波兰语成分重排。
  • 通过对齐源语言和目标语言的语篇结构,提升英波机器翻译的流畅性和交际恰当性。

提出的方法

  • 将中心理论扩展,以包含定指性、指示代词、所有格修饰语以及所有名词短语的可分级中心值。
  • 基于多种线索分配中心值:代词指代、重复提及、定指性、语法功能和句法突出性。
  • 采用三级处理流水线:预处理表(用于回指和特殊结构)、偏好表(用于一般排序偏好)和冲突分辨表(用于解决规则冲突)。
  • 当规则存在歧义时,利用语料库数据中的统计偏好(例如,VSO 66%,OVS 50%)指导排序。
  • 结合指代距离和成分长度以优化排序决策。
  • 将中心理论与功能句子主位(FSP)结合,以预测波兰语中的信息结构。

实验结果

研究问题

  • RQ1如何在英语中建模语篇级信息结构,以支持向波兰语的准确翻译?
  • RQ2哪些语言线索能可靠地预测波兰语中的语篇突出性和成分顺序?
  • RQ3如何扩展中心理论以处理非主语成分和分级显著性?
  • RQ4定指性、代词指代和句法功能在决定波兰语成分顺序中起什么作用?
  • RQ5如何将统计规律与句法线索结合,以生成稳健且上下文敏感的波兰语词序?

主要发现

  • 扩展的中心模型成功预测了英语中的语篇突出性,其中代词指代和重复提及是中心的主要指标。
  • 定指性和指示性修饰语显著提升了中心值,尤其在与先前提及结合时效果更明显。
  • 当结合中心值与统计偏好时,波兰语成分顺序的预测最为准确(例如,VSO结构频率达66%)。
  • 预处理表通过应用风格规则和焦点绑定规则,解决了诸如“we”或“only”等模糊情况。
  • 冲突分辨表通过优先将较长成分置于句末(例如,当长度(O) ≥ 长度(S)时,OVS结构的成功率为100%),有效解决了规则间的冲突,提升了准确性。
  • 通过将源语言语篇结构与目标语言词序对齐,该系统显著提升了交际恰当性,尤其在包含状语和话题化成分的复杂从句中表现突出。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。