[Paper Review] Unsupervised Translation of Programming Languages
The paper trains TransCoder, a fully unsupervised neural transcompiler, to translate functions among C++, Java, and Python using monolingual code, and releases a 852-function parallel test set with unit tests showing strong performance over rule-based baselines.
A transcompiler, also known as source-to-source translator, is a system that converts source code from a high-level programming language (such as C++ or Python) to another. Transcompilers are primarily used for interoperability, and to port codebases written in an obsolete or deprecated language (e.g. COBOL, Python 2) to a modern one. They typically rely on handcrafted rewrite rules, applied to the source code abstract syntax tree. Unfortunately, the resulting translations often lack readability, fail to respect the target language conventions, and require manual modifications in order to work properly. The overall translation process is timeconsuming and requires expertise in both the source and target languages, making code-translation projects expensive. Although neural models significantly outperform their rule-based counterparts in the context of natural language translation, their applications to transcompilation have been limited due to the scarcity of parallel data in this domain. In this paper, we propose to leverage recent approaches in unsupervised machine translation to train a fully unsupervised neural transcompiler. We train our model on source code from open source GitHub projects, and show that it can translate functions between C++, Java, and Python with high accuracy. Our method relies exclusively on monolingual source code, requires no expertise in the source or target languages, and can easily be generalized to other programming languages. We also build and release a test set composed of 852 parallel functions, along with unit tests to check the correctness of translations. We show that our model outperforms rule-based commercial baselines by a significant margin.
Motivation & Objective
- Motivate automated translation of existing codebases across languages without parallel data or expert rules.
- Develop a fully unsupervised transcompiler trained on GitHub code using cross-lingual pretraining, denoising auto-encoding, and back-translation.
- Demonstrate that the unsupervised model can outperform rule-based and commercial baselines on function-level translations.
- Provide a validation/test set of parallel functions with unit tests to evaluate translation correctness.
Proposed method
- Use a single Transformer-based encoder-decoder model shared across C++, Java, and Python.
- Pretrain with cross-lingual masked language modeling (XLM) on monolingual code for cross-language representations.
- Apply denoising auto-encoding to make the decoder generate valid sequences and robust representations.
- Leverage back-translation to create pseudo-parallel data between language pairs.
- Evaluate translations using computational accuracy via unit tests, alongside reference match and BLEU metrics.
Experimental results
Research questions
- RQ1Can a fully unsupervised neural transcompiler learn to translate between C++, Java, and Python using only monolingual code?
- RQ2How does TransCoder perform compared to rule-based and commercial baselines on function-level translations?
- RQ3What evaluation metrics best reflect functional correctness beyond BLEU or reference overlap?
- RQ4What is the impact of beam search and input preprocessing on translation quality?
- RQ5Can the approach generalize to other programming languages beyond the three studied?
Key findings
- TransCoder achieves higher computational accuracy than baselines across language directions (e.g., C++→Java: 60.91%; Java→Python: 34.99%).
- Reference match and BLEU do not correlate well with actual functional correctness, as many correct translations do not match references exactly or have high BLEU.
- Beam search improves computational accuracy substantially (up to 33.7% in some directions) when using unit-test validation.
- TransCoder outperforms a Java→Python baseline and a C++→Java commercial baseline in computational accuracy.
- Keeping source code comments increases anchor points and improves cross-language alignment, boosting performance.
- The model learns to map language-specific constructs and standard library usage across languages without supervision.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.