[論文レビュー] A systematic literature review on source code similarity measurement and clone detection: techniques, applications, and challenges
本システマティックレビューでは、136件の一次研究を分析し、8つの手法と5つの分野におけるソースコード類似度測定およびクローン検出の技術、ツール、応用、課題を評価した。80のツールが同定され、Java(49%)とC/C++(37%)への強い対応が確認された。一方、公開可能なデータセットはわずか8つにとどまり、実証的検証の欠如、ハイブリッド手法の不足、マルチパラダイム言語対応の不十分さといった主な課題が明らかになった。
Measuring and evaluating source code similarity is a fundamental software engineering activity that embraces a broad range of applications, including but not limited to code recommendation, duplicate code, plagiarism, malware, and smell detection. This paper proposes a systematic literature review and meta-analysis on code similarity measurement and evaluation techniques to shed light on the existing approaches and their characteristics in different applications. We initially found over 10000 articles by querying four digital libraries and ended up with 136 primary studies in the field. The studies were classified according to their methodology, programming languages, datasets, tools, and applications. A deep investigation reveals 80 software tools, working with eight different techniques on five application domains. Nearly 49% of the tools work on Java programs and 37% support C and C++, while there is no support for many programming languages. A noteworthy point was the existence of 12 datasets related to source code similarity measurement and duplicate codes, of which only eight datasets were publicly accessible. The lack of reliable datasets, empirical evaluations, hybrid methods, and focuses on multi-paradigm languages are the main challenges in the field. Emerging applications of code similarity measurement concentrate on the development phase in addition to the maintenance.
研究の動機と目的
- ソースコード類似度測定およびクローン検出に用いられる既存の技術を特定・分類すること。
- 研究間で使用されたツール、プログラミング言語、データセットの分布を分析すること。
- 分野における実証的評価の成熟度と信頼性を評価すること。
- 保守フェーズを超えた、とりわけ開発フェーズにおける新たな応用を同定すること。
- 公開データセットの不足、マルチパラダイム対応の限定、ハイブリッド手法の不十分さといった重要な課題を強調すること。
提案手法
- 関連する研究を同定するために、4つのデジタル図書館を用いてシステマティックレビューを実施した。
- 事前に定めた包含・除外基準に基づき、136件の一次研究をスクリーニング・選定した。
- 研究を手法、プログラミング言語、データセット、ツール、応用分野ごとに分類した。
- ツールの特徴と技術の普及度に注目して、研究間の合成を図るメタアナリシスを実施した。
- 80のソフトウェアツールを8つの異なる検出手法と5つの応用分野にマッピングした。
- データセットの可用性と品質を評価し、12のデータセットのうち8つが公開可能であることを同定した。
実験結果
リサーチクエスチョン
- RQ1ソースコード類似度測定およびクローン検出で主に用いられる技術は何か?
- RQ2既存のツールや研究が主に標的とするプログラミング言語は何か?
- RQ3コード類似度およびクローン検出に用いられる公開可能なデータセットはいくつ存在し、その品質はどの程度か?
- RQ4ソフトウェア工学におけるコード類似度測定の主な応用は何か?
- RQ5この研究分野の進展を妨げる主な課題は何か?
主な発見
- 80のソフトウェアツールが同定され、そのうち49%がJavaを対象としており、37%がCおよびC++を対象としていた。
- コード類似度およびクローン検出に関連する12のデータセットのうち、公開可能なのは8つにとどまった。
- 大多数の研究が、クローン検出やバグ検出といった保守フェーズの応用に焦点を当てていた。
- 研究間で実証的評価やベンチマークの欠如が顕著で、再現性の制限要因となっている。
- 精度向上の可能性を秘めているにもかかわらず、複数の技術を組み合わせたハイブリッド手法は未だ十分に検討されていない。
- マルチパラダイム言語やあまり使われないプログラミング言語への対応は、依然として最小限または存在しない。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。