[논문 리뷰] A Comprehensive Survey on Pretrained Foundation Models: A History from BERT to ChatGPT
본 고찰은 NLP, CV 및 Graph Learning 전반에 걸친 Pretrained Foundation Models (PFMs)를 검토하고, BERT에서 ChatGPT로의 진화를 추적하며 아키텍처, pretraining tasks, 통합 PFMs, 효율성, 보안 및 향후 과제를 논의한다.
Pretrained Foundation Models (PFMs) are regarded as the foundation for various downstream tasks with different data modalities. A PFM (e.g., BERT, ChatGPT, and GPT-4) is trained on large-scale data which provides a reasonable parameter initialization for a wide range of downstream applications. BERT learns bidirectional encoder representations from Transformers, which are trained on large datasets as contextual language models. Similarly, the generative pretrained transformer (GPT) method employs Transformers as the feature extractor and is trained using an autoregressive paradigm on large datasets. Recently, ChatGPT shows promising success on large language models, which applies an autoregressive language model with zero shot or few shot prompting. The remarkable achievements of PFM have brought significant breakthroughs to various fields of AI. Numerous studies have proposed different methods, raising the demand for an updated survey. This study provides a comprehensive review of recent research advancements, challenges, and opportunities for PFMs in text, image, graph, as well as other data modalities. The review covers the basic components and existing pretraining methods used in natural language processing, computer vision, and graph learning. Additionally, it explores advanced PFMs used for different data modalities and unified PFMs that consider data quality and quantity. The review also discusses research related to the fundamentals of PFMs, such as model efficiency and compression, security, and privacy. Finally, the study provides key implications, future research directions, challenges, and open problems in the field of PFMs. Overall, this survey aims to shed light on the research of the PFMs on scalability, security, logical reasoning ability, cross-domain learning ability, and the user-friendly interactive ability for artificial general intelligence.
연구 동기 및 목표
- NLP, GL 전반에 걸친 PFMs의 개발을 조사하고 BERT에서 ChatGPT까지의 역사를 추적한다.
- 다양한 모달리티에 걸친 기본 구성요소, 학습 패러다임 및 사전학습 과제를 분석한다.
- 통합 PFMs, 모델 효율성, 압축, 보안 및 프라이버시와 같은 고급 주제를 논의한다.
- PFMs의 향후 연구를 안내하기 위한 도전과제 및 미해결 문제를 식별한다.
- 부록을 통해 PFMs와 관련된 평가 지표 및 데이터셋에 대한 지침을 제공한다.
제안 방법
- 트랜스포머 기반 PFMs를 지배적 아키텍처로 설명한다.
- 학습 메커니즘(감독학습, 반지도학습, SSL, RL)을 분류하고 사전학습에서의 역할을 설명한다.
- NLP 사전학습 과제(MLM, DAE, RTD, NSP, SOP)와 CV/GL SSL 전략을 요약한다.
- Autoregressive, contextual, permuted LMs와 같은 모델 설계 선택과 주요 모델(GPT, BERT, XLNet, MPNet)을 개요한다.
- 지시 정렬 방법(RLHF, chain-of-thought 등)을 사용하여 출력이 인간의 선호도와 일치하도록 논의한다.
- 통합 PFMs, 효율성, 압축, 보안 및 프라이버시와 같은 고급 주제에 대한 논의를 종합한다.
실험 결과
연구 질문
- RQ1NLP, CV, 및 GL 전반에서 PFMs를 가능하게 하는 핵심 구성 요소와 학습 메커니즘은 무엇인가?
- RQ2표현 학습 및 다운스트림 성능을 향상시키기 위해 모달리티 전반에서 사전학습 과제가 어떻게 진화해 왔는가?
- RQ3Autoregressive, contextual, 및 permuted language models 간의 설계 트레이드오프는 무엇인가?
- RQ4다중모달 데이터에 걸친 통합 PFMs의 현황은 무엇이며 남아 있는 도전과제는 무엇인가?
- RQ5PFMs의 효율성, 보안 및 프라이버시에 관한 미해결 과제와 향후 방향은 무엇인가?
주요 결과
- 트랜스포머는 NLP, CV, 및 GL 전반에서 확장 가능한 PFMs를 가능하게 하는 핵심 아키텍처로 남아 있다.
- 사전학습 과제와 학습 메커니즘(SSL, RL, 및 감독 신호)은 강건한 특징 표현과 신속한 미세조정을 뒷받침한다.
- NLP 사전학습 과제에는 MLM, DAE, RTD, NSP, 및 SOP가 포함되며, XLNet 및 MPNet과 같은 확장은 permuted/hybrid 목적을 도입한다.
- 텍스트, 이미지, 오디오를 처리할 수 있는 통합 PFMs가 등장하고 있으며, GPT-4 등으로 예시된다.
- 고급 주제는 모델 효율성, 압축, 보안 및 프라이버시를 강조하며 실제 배포 및 거버넌스 문제를 다룬다.
- 본 연구는 확장성, 교차 도메인 학습, 추론 및 대화형 AI 능력의 향후 연구 방향을 강조한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.