Skip to main content
QUICK REVIEW

[論文レビュー] A Survey of Neural Code Intelligence: Paradigms, Advances and Beyond

Qiushi Sun, Zhirui Chen|arXiv (Cornell University)|Mar 21, 2024
Machine Learning in Bioinformatics被引用数 5
ひとこと要約

この調査は、神経コード知能の進化を3つのフェーズ—コードの神経言語モデル、CodePTMs、CodeLLMs—を横断して系統的にレビューし、タスク、ベンチマーク、分野横断の相乗効果、将来の方向性を扱う。

ABSTRACT

Neural Code Intelligence -- leveraging deep learning to understand, generate, and optimize code -- holds immense potential for transformative impacts on the whole society. Bridging the gap between Natural Language and Programming Language, this domain has drawn significant attention from researchers in both research communities over the past few years. This survey presents a systematic and chronological review of the advancements in code intelligence, encompassing over 50 representative models and their variants, more than 20 categories of tasks, and an extensive coverage of over 680 related works. We follow the historical progression to trace the paradigm shifts across different research phases (e.g., from modeling code with recurrent neural networks to the era of Large Language Models). Concurrently, we highlight the major technical transitions in models, tasks, and evaluations spanning through different stages. For applications, we also observe a co-evolving shift. It spans from initial endeavors to tackling specific scenarios, through exploring a diverse array of tasks during its rapid expansion, to currently focusing on tackling increasingly complex and varied real-world challenges. Building on our examination of the developmental trajectories, we further investigate the emerging synergies between code intelligence and broader machine intelligence, uncovering new cross-domain opportunities and illustrating the substantial influence of code intelligence across various domains. Finally, we delve into both the opportunities and challenges associated with this field, alongside elucidating our insights on the most promising research directions. An ongoing, dynamically updated project and resources associated with this survey have been released at https://github.com/QiushiSun/Awesome-Code-Intelligence.

研究の動機と目的

  • 神経コード知能の歴史的進化と各フェーズにおけるパラダイムシフトを追跡する。
  • コード関連のタスクとベンチマークを、構造化された分析のために一貫したカテゴリに分類する。
  • コアとなるモデルアーキテクチャ、学習目的、コード構造(AST、データフロー、制御フロー)の役割を要約する。
  • 応用事例、分野横断の相乗効果、実務上の課題を議論し、今後の研究方針を導くために。

提案手法

  • 50を超える代表的なモデルと680以上の関連研究を対象とした体系的・時系列的レビューを実施する。
  • モデルを、コードの神経言語モデル、コード事前学習モデル(CodePTMs)、およびCodeLLMsの3つの時代に整理する。
  • 構造的なコード情報を重視して、アーキテクチャの選択、学習データ、目的を分析する。
  • 20以上のカテゴリにわたるコード関連タスクとベンチマークを網羅的に整理・要約する。
  • 分野横断的な統合と、実世界のアプリケーションと評価への影響を議論する。
  • 未解決の課題と今後の有望な方向性を特定する。
Figure 1 : Cumulative number of publications/preprints related to neural code intelligence (from arXiv). Over the past few years, the number of articles has been steadily increasing.
Figure 1 : Cumulative number of publications/preprints related to neural code intelligence (from arXiv). Over the past few years, the number of articles has been steadily increasing.

実験結果

リサーチクエスチョン

  • RQ1初期の神経LMからCodePTMsおよびCodeLLMsへ至る神経コード知能の主要なパラダイムシフトは何か?
  • RQ2CodePTMsはパフォーマンス、データ要件、タスク網羅性の点で大規模言語モデルとどのように比較できるか?
  • RQ3コード知能の進展を牽引する主要なベンチマークとタスクは何で、時代を超えてどのように進化するか?
  • RQ4コード知能をより広い機械知能や実世界のアプリケーションと統合する際の分野横断的な機会と課題は何か?

主な発見

  • コードの神経言語モデル、コードの事前学習モデル、そしてCodeLLMsという3つの段階的なフェーズが進展を推進してきた。
  • CodeBERTやCodeT5のようなCodePTMsが、コードの事前学習+ファインチューニングのパラダイムを確立した。
  • CodeLLMsはプロンプト利用やイン-context学習への移行を示し、コードのみのタスクを超えて実世界のシナリオへ拡大している。
  • 20以上のカテゴリにわたる多数のタスクと数千のデータセットという幅広いベンチマークが分野の進展を支えている。
  • 分野横断的な相乗効果と実世界の課題が、コード知能の将来の方向性を形作っている。
Figure 2 : A chronological overview of representative works in neural code intelligence over recent years. Works are differentiated by background colors to represent distinct evolutionary phases: ${\color[rgb]{1,0.88671875,0.58203125}\definecolor[named]{pgfstrokecolor}{rgb}{1,0.88671875,0.58203125}\
Figure 2 : A chronological overview of representative works in neural code intelligence over recent years. Works are differentiated by background colors to represent distinct evolutionary phases: ${\color[rgb]{1,0.88671875,0.58203125}\definecolor[named]{pgfstrokecolor}{rgb}{1,0.88671875,0.58203125}\

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。