Skip to main content
QUICK REVIEW

[논문 리뷰] A Survey of Deep Learning Techniques for Neural Machine Translation

Shuoheng Yang, Yuxin Wang|arXiv (Cornell University)|2020. 02. 18.
Natural Language Processing Techniques참고 문헌 133인용 수 98
한 줄 요약

This paper surveys the origin and development of Neural Machine Translation (NMT), categorizing models by architecture and training approaches, highlighting attention and coverage mechanisms, and outlining future research trends.

ABSTRACT

In recent years, natural language processing (NLP) has got great development with deep learning techniques. In the sub-field of machine translation, a new approach named Neural Machine Translation (NMT) has emerged and got massive attention from both academia and industry. However, with a significant number of researches proposed in the past several years, there is little work in investigating the development process of this new technology trend. This literature survey traces back the origin and principal development timeline of NMT, investigates the important branches, categorizes different research orientations, and discusses some future research trends in this field.

연구 동기 및 목표

  • Trace the origin and development timeline of Neural Machine Translation (NMT).
  • Categorize NMT models by architectural orientation and training approaches.
  • Discuss key components (attention, vocabulary coverage) and their impact on translation quality.
  • Evaluate advanced NMT models and potential future directions.

제안 방법

  • Literature review of historical MT paradigms (rule-based, statistical, neural).
  • Classification of NMT variants by architecture (RNN-based, CNN-based, Transformer).
  • Discussion of encoder–decoder structure and training/inference procedures.
  • Analysis of attention mechanisms and coverage problems in NMT.
  • Overview of advanced models and practical implications for speed and quality.

실험 결과

연구 질문

  • RQ1What are the major development stages and model families in NMT history?
  • RQ2How do attention and coverage mechanisms improve translation quality in NMT?
  • RQ3What are the trade-offs between RNN, CNN, and Transformer-based NMT models?
  • RQ4What are the highlighted future trends and directions in NMT research?

주요 결과

  • NMT evolved from shallow to deep architectures and to attention-based and Transformer models.
  • Encoder–decoder is the core structure for end-to-end NMT and is enhanced via depth, bidirectionality, and specialized units (LSTM/GRU).
  • Attention mechanisms address long-range dependencies and improve translation quality beyond early RNN/NLM baselines.
  • CNN-based NMT offers speed advantages but historically struggled with long-range dependencies until attention/Transformer refinements.
  • Vocabulary coverage and related mechanisms are essential considerations in current NMT designs.
  • The survey discusses advanced models (e.g., GNMT, MoE) and anticipates future research directions in NMT.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.