[논문 리뷰] Large Language Models for Software Engineering: A Systematic Literature Review
대형 언어 모델을 소프트웨어 엔지니어링에 적용한 229편의 논문(2017–2023)을 분석하고, 모델 유형, 데이터 관행, 최적화/평가 전략, SE 과제를 분류한 체계적 문헌 고찰.
Large Language Models (LLMs) have significantly impacted numerous domains, including Software Engineering (SE). Many recent publications have explored LLMs applied to various SE tasks. Nevertheless, a comprehensive understanding of the application, effects, and possible limitations of LLMs on SE is still in its early stages. To bridge this gap, we conducted a systematic literature review (SLR) on LLM4SE, with a particular focus on understanding how LLMs can be exploited to optimize processes and outcomes. We select and analyze 395 research papers from January 2017 to January 2024 to answer four key research questions (RQs). In RQ1, we categorize different LLMs that have been employed in SE tasks, characterizing their distinctive features and uses. In RQ2, we analyze the methods used in data collection, preprocessing, and application, highlighting the role of well-curated datasets for successful LLM for SE implementation. RQ3 investigates the strategies employed to optimize and evaluate the performance of LLMs in SE. Finally, RQ4 examines the specific SE tasks where LLMs have shown success to date, illustrating their practical contributions to the field. From the answers to these RQs, we discuss the current state-of-the-art and trends, identifying gaps in existing research, and flagging promising areas for future study. Our artifacts are publicly available at https://github.com/xinyi-hou/LLM4SE_SLR.
연구 동기 및 목표
- SE 과제에 사용된 LLM이 어떤 것이며, 아키텍처 및 특징에 따라 어떻게 분류되는지 맵핑한다.
- LLM4SE 연구의 데이터 수집, 전처리 및 표현 방식 분석.
- SE에서 LLM을 위해 사용된 최적화 및 평가 전략 식별.
- LLMs가 효과를 보인 SE 과제를 식별하고 경향, 격차 및 향후 방향을 도출한다.
제안 방법
- Kitchenham 스타일의 체계적 문헌고찰 프로세스(계획, 수행, 분석)를 따랐다.
- 수동으로 식별된 관련 논문들로부터 준 골드 스탠다드를 구축한 뒤, 포괄성을 확보하기 위해 자동 검색 및 스노볼링을 실시했다.
- 명시적 포함/제외 기준과 10항 품질 평가 체크리스트를 적용하여 고품질의 주요 연구를 선정했다.
- SE 과제 카테고리, LLM 카테고리, 데이터 처리, 최적화 알고리즘, 평가 지표 및 SE 활동에 대한 데이터 추출을 수행했다.
- 발표 장소, 연도, 아키텍처(인코더-전용, 인코더-디코더, 디코더-전용)에 걸친 기술 분석과 트렌드 분석을 수행했다.
- 최신 연구 동향, 도전과제 및 향후 연구 방향을 정리하기 위해 결과를 합성했다.
실험 결과
연구 질문
- RQ1RQ1: 지금까지 SE 과제를 해결하기 위해 어떤 LLM이 고용되었는가?
- RQ2RQ2: SE 관련 데이터세트가 어떻게 수집되고, 전처리되며, LLM에서 어떻게 사용되는가?
- RQ3RQ3: LLM4SE를 최적화하고 평가하는 데 어떤 기법이 사용되는가?
- RQ4RQ4: 지금까지 LLM4SE를 사용해 효과적으로 다뤄진 SE 과제는 무엇인가?
주요 결과
- 본 연구는 2017–2023년의 229편의 논문을 분석한 SE를 위한 LLM 기반 솔루션에 관한 최초의 포괄적 SLR이다.
- 수집된 문헌에서 SE 과제에 사용된 LLM은 50개가 넘는다.
- 인코더-전용, 인코더-디코더, 디코더-전용 LLM이 사용되며, 2023년에 디코더-전용의 우세가 나타난다.
- 문헌은 SE 과제에 맞춘 다양한 데이터 처리 관행과 최적화 및 평가 접근법의 다양성을 보여준다.
- SE 과자는 55개의 서로 다른 활동에 걸치며, 6대 핵심 SE 활동(요구사항, 설계, 개발, 품질 보증, 유지보수, 관리)으로 묶인다.
- 빠른 성장 추세가 뚜렷하며, 2022년에서 2023년 초에 급격히 증가했고, arXiv에 발표 비중이 크게 나타난다(지속적인 급속 발전을 반영).
- 리뷰는 모델 선택, 데이터 처리, 미세조정, 평가 및 배포 고려사항을 포함한 LLM4SE의 도전과제와 향후 연구 방향을 제시한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.