[Paper Review] A Comprehensive Survey on Pretrained Foundation Models: A History from BERT to ChatGPT
This survey reviews Pretrained Foundation Models (PFMs) across NLP, CV, and Graph Learning, tracing the evolution from BERT to ChatGPT and discussing architectures, pretraining tasks, unified PFMs, efficiency, security, and future challenges.
Pretrained Foundation Models (PFMs) are regarded as the foundation for various downstream tasks with different data modalities. A PFM (e.g., BERT, ChatGPT, and GPT-4) is trained on large-scale data which provides a reasonable parameter initialization for a wide range of downstream applications. BERT learns bidirectional encoder representations from Transformers, which are trained on large datasets as contextual language models. Similarly, the generative pretrained transformer (GPT) method employs Transformers as the feature extractor and is trained using an autoregressive paradigm on large datasets. Recently, ChatGPT shows promising success on large language models, which applies an autoregressive language model with zero shot or few shot prompting. The remarkable achievements of PFM have brought significant breakthroughs to various fields of AI. Numerous studies have proposed different methods, raising the demand for an updated survey. This study provides a comprehensive review of recent research advancements, challenges, and opportunities for PFMs in text, image, graph, as well as other data modalities. The review covers the basic components and existing pretraining methods used in natural language processing, computer vision, and graph learning. Additionally, it explores advanced PFMs used for different data modalities and unified PFMs that consider data quality and quantity. The review also discusses research related to the fundamentals of PFMs, such as model efficiency and compression, security, and privacy. Finally, the study provides key implications, future research directions, challenges, and open problems in the field of PFMs. Overall, this survey aims to shed light on the research of the PFMs on scalability, security, logical reasoning ability, cross-domain learning ability, and the user-friendly interactive ability for artificial general intelligence.
Motivation & Objective
- Survey the development of PFMs across NLP, CV, and GL and trace their history from BERT to ChatGPT.
- Analyze basic components, training paradigms, and pretraining tasks across modalities.
- Discuss advanced topics such as unified PFMs, model efficiency, compression, security, and privacy.
- Identify challenges and open problems to guide future research in PFMs.
- Provide guidance on evaluation metrics and datasets relevant to PFMs through appendices.
Proposed method
- Describe Transformer-based PFMs as the dominant architecture.
- Categorize learning mechanisms (supervised, semi-supervised, SSL, RL) and their roles in pretraining.
- Summarize NLP pretraining tasks (MLM, DAE, RTD, NSP, SOP) and CV/GL SSL strategies.
- Outline model design choices (autoregressive, contextual, permuted LMs) and key models (e.g., GPT, BERT, XLNet, MPNet).
- Discuss instruction-aligning methods (e.g., RLHF, chain-of-thought) used to align outputs with human preferences.
- Synthesize discussion on advanced topics like unified PFMs, efficiency, compression, security, and privacy.
Experimental results
Research questions
- RQ1What are the core components and learning mechanisms enabling PFMs across NLP, CV, and GL?
- RQ2How have pretraining tasks evolved across modalities to improve representation learning and downstream performance?
- RQ3What are the design tradeoffs between autoregressive, contextual, and permuted language models?
- RQ4What is the state of unified PFMs across multimodal data, and what challenges remain?
- RQ5What are the open challenges and future directions in efficiency, security, and privacy for PFMs?
Key findings
- Transformers remain the central architecture enabling scalable PFMs across NLP, CV, and GL.
- Pretraining tasks and learning mechanisms (SSL, RL, and supervised signals) underpin robust feature representations and rapid fine-tuning.
- NLP pretraining tasks include MLM, DAE, RTD, NSP, and SOP, with extensions like XLNet and MPNet introducing permuted/hybrid objectives.
- Emerging trends include unified PFMs capable of handling text, images, and audio, exemplified by models like GPT-4 and others.
- Advanced topics emphasize model efficiency, compression, security, and privacy, addressing practical deployment and governance concerns.
- The survey highlights future research directions in scalability, cross-domain learning, reasoning, and interactive AI capabilities.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.