Skip to main content
QUICK REVIEW

[논문 리뷰] A Comprehensive Overview of Large Language Models

Humza Naveed, Asad Ullah Khan|arXiv (Cornell University)|2023. 07. 12.
Topic Modeling인용 수 351
한 줄 요약

이 논문은 대형 언어 모델(LLMs)에 대한 독자적이고 포괄적인 조사로 구성되어 있으며, 아키텍처, 학습, 미세조정, 다중 모달 확장, 데이터셋, 평가, 효율성 및 향후 과제를 다룬다. 또한 유명한 사전 학습된 LLM의 상세 요약과 연구자 및 실무자를 위한 실용적 지침을 제공한다.

ABSTRACT

Large Language Models (LLMs) have recently demonstrated remarkable capabilities in natural language processing tasks and beyond. This success of LLMs has led to a large influx of research contributions in this direction. These works encompass diverse topics such as architectural innovations, better training strategies, context length improvements, fine-tuning, multi-modal LLMs, robotics, datasets, benchmarking, efficiency, and more. With the rapid development of techniques and regular breakthroughs in LLM research, it has become considerably challenging to perceive the bigger picture of the advances in this direction. Considering the rapidly emerging plethora of literature on LLMs, it is imperative that the research community is able to benefit from a concise yet comprehensive overview of the recent developments in this field. This article provides an overview of the existing literature on a broad range of LLM-related concepts. Our self-contained comprehensive overview of LLMs discusses relevant background concepts along with covering the advanced topics at the frontier of research in LLMs. This review article is intended to not only provide a systematic survey but also a quick comprehensive reference for the researchers and practitioners to draw insights from extensive informative summaries of the existing works to advance the LLM research.

연구 동기 및 목표

  • 대형 언어 모델(LLMs)의 최근 발전에 대한 간결하고 포괄적인 개요를 제공한다.
  • 정교한 정보로 사전 학습된 LLM의 아키텍처 및 학습 세부 정보를 요약한다.
  • 미세 조정, 다중 모달 LLM, 보강 LLM, 데이터셋, 벤치마크, 평가 및 배포 고려사항을 논의한다.

제안 방법

  • 배경, 아키텍처, 학습 파이프라인 및 전략을 제시하기 위해 LLM 문헌을 조사한다.
  • 주요 사전 학습 LLM을 표로 요약한다.
  • 구성, 평가, 데이터셋, 벤치마크 및 실무자를 위한 실용적 고려사항을 논의한다.
Figure 1: The trend of papers released over years containing keywords “Large Language Model”, “Large Language Model + Fine-Tuning”, and “Large Language Model + Alignment”.
Figure 1: The trend of papers released over years containing keywords “Large Language Model”, “Large Language Model + Fine-Tuning”, and “Large Language Model + Alignment”.

실험 결과

연구 질문

  • RQ1주요 LLM들에서 핵심 아키텍처 선택과 학습 전략은 무엇인가?
  • RQ2미세 조정, 지시 학습(instruction-tuning), 정렬 조정(alignment-tuning)이 제로샷 및 파샷 성능에 어떤 영향을 미치는가?
  • RQ3LLM을 평가하는 데 사용되는 데이터셋, 벤치마크 및 평가 방법은 무엇이며 확인된 문제점은 무엇인가?
  • RQ4LLM 연구 및 실무에서의 효율성, 배포 및 안전성 고려사항은 무엇인가?

주요 결과

  • LLMs는 점차 지시 학습에 맞춰 조정되고 있으며 점점 더 오픈 소스 모델로 진화해 왔다.
  • 추론 및 컨텍스트 내 학습과 같은 새로운 능력이 대규모에서 나타나 광범위한 활용에 영향을 준다.
  • 비용 감소를 위해 매개변수 효율적 튜닝, 가지치기, 양자화, MoE, 컨텍스트 길이 전략 등의 효율성 접근법이 적극 연구되고 있다.
  • LLM 평가에 광범위한 데이터셋과 벤치마크가 사용되며, 사실적 정확성과 인간 선호에의 정렬에 중점을 둔다.
  • 연구자들은 로봇공학 및 도구 사용을 포함한 다중 모달 및 에이전트 지향 설정으로 LLM을 확장하고 있다.
  • 과제에는 사실적 정확성, 인간 가치와의 정렬, 안전성, 그리고 자원 집중적 학습 및 추론이 포함된다.
Figure 2: Chronological display of LLM releases: light blue rectangles represent ‘pre-trained’ models, while dark rectangles correspond to ‘instruction-tuned’ models. Models on the upper half signify open-source availability, whereas those on the bottom half are closed-source. The chart illustrates
Figure 2: Chronological display of LLM releases: light blue rectangles represent ‘pre-trained’ models, while dark rectangles correspond to ‘instruction-tuned’ models. Models on the upper half signify open-source availability, whereas those on the bottom half are closed-source. The chart illustrates

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.