Skip to main content
QUICK REVIEW

[論文レビュー] Large Language Models for Software Engineering: A Systematic Literature Review

Xinyi Hou, Yanjie Zhao|arXiv (Cornell University)|Aug 21, 2023
Software Engineering Research被引用数 111
ひとこと要約

システマティック文献レビューは、ソフトウェア工学への大規模言語モデル適用に関する229件の論文(2017–2023)を分析し、モデル、データ実践、最適化/評価戦略、SEタスクを分類します。

ABSTRACT

Large Language Models (LLMs) have significantly impacted numerous domains, including Software Engineering (SE). Many recent publications have explored LLMs applied to various SE tasks. Nevertheless, a comprehensive understanding of the application, effects, and possible limitations of LLMs on SE is still in its early stages. To bridge this gap, we conducted a systematic literature review (SLR) on LLM4SE, with a particular focus on understanding how LLMs can be exploited to optimize processes and outcomes. We select and analyze 395 research papers from January 2017 to January 2024 to answer four key research questions (RQs). In RQ1, we categorize different LLMs that have been employed in SE tasks, characterizing their distinctive features and uses. In RQ2, we analyze the methods used in data collection, preprocessing, and application, highlighting the role of well-curated datasets for successful LLM for SE implementation. RQ3 investigates the strategies employed to optimize and evaluate the performance of LLMs in SE. Finally, RQ4 examines the specific SE tasks where LLMs have shown success to date, illustrating their practical contributions to the field. From the answers to these RQs, we discuss the current state-of-the-art and trends, identifying gaps in existing research, and flagging promising areas for future study. Our artifacts are publicly available at https://github.com/xinyi-hou/LLM4SE_SLR.

研究の動機と目的

  • どのLLMsがSEタスクに用いられてきたかを、アーキテクチャと特徴の観点で分類してマッピングする。
  • LLM4SE研究のデータ収集、前処理、表現方法を分析する。
  • SEにおけるLLMsの最適化および評価戦略を特定する。
  • LLMsが効果を示したSEタスクを特定し、動向・ギャップ・将来の方向性を導出する。

提案手法

  • Kitchenham風の系統的文献レビュー手順(計画、実施、分析)を採用した。
  • manually identified relevant papers から準ゴールドスタンダードを構築し、その後自動検索とスノーボール法を用いて網羅性を高めた。
  • 明示的な包含/排除基準と10項目の品質評価チェックリストを適用して高品質な一次研究を選択した。
  • SEタスクのカテゴリー、LLMカテゴリー、データ処理、最適化アルゴリズム、評価指標、SE活動に関するデータ抽出を実施した。
  • 公表会議・ジャーナル・プラットフォーム別の記述分析と、年次・アーキテクチャ(エンコーダ専用、エンコーダ-デコーダ、デコーダ専用)ごとのトレンド分析を実施した。
  • 結果を総括して最先端の状況、課題、将来の研究方向性をマッピングした。

実験結果

リサーチクエスチョン

  • RQ1RQ1: これまでにSEタスクを解決するために用いられてきたLLMはどれか?
  • RQ2RQ2: SE関連データセットはどのように収集・前処理・利用されているか?
  • RQ3RQ3: LLM4SEを最適化・評価するためにどの技術が用いられているか?
  • RQ4RQ4: 現時点でLLM4SEを用いて効果的に扱われたSEタスクは何か?

主な発見

  • 本研究は、SEのためのLLMベース解決策に関する最初の包括的なSLRであり、2017–2023年の229論文を分析した。
  • 収集された文献にはSEタスクのために50以上の異なるLLMが使用されていた。
  • エンコーダ専用、エンコーダ-デコーダ、デコーダ専用のLLMが使用されており、2023年にはデコーダ専用の優位性が顕著に出現した。
  • 文献には多様なデータ処理実践と、SEタスクに合わせたさまざまな最適化・評価手法が示されている。
  • SEタスクは55の異なる活動を含み、6つのコアSE活動(要件、設計、開発、QA、保守、管理)に整理されている。
  • 公開の急速な成長傾向が見られ、2022年〜2023年前半に公表が急増し、arXiv掲載の論文が substantial share を占めている(継続的な急速発展を反映)。
  • レビューは課題を特定し、LLM4SEの将来研究方向として、モデル選択、データ処理、ファインチューニング、評価、デプロイメントの考慮事項を含む方向性を提案する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。