[论文解读] A Comprehensive Survey on Pretrained Foundation Models: A History from BERT to ChatGPT
本综述系统评估跨自然语言处理、计算机视觉与图形学习的预训练基础模型(PFMs),追踪从 BERT 到 ChatGPT 的演变,并讨论架构、预训练任务、统一 PFMs、效率、安全性以及未来挑战。
Pretrained Foundation Models (PFMs) are regarded as the foundation for various downstream tasks with different data modalities. A PFM (e.g., BERT, ChatGPT, and GPT-4) is trained on large-scale data which provides a reasonable parameter initialization for a wide range of downstream applications. BERT learns bidirectional encoder representations from Transformers, which are trained on large datasets as contextual language models. Similarly, the generative pretrained transformer (GPT) method employs Transformers as the feature extractor and is trained using an autoregressive paradigm on large datasets. Recently, ChatGPT shows promising success on large language models, which applies an autoregressive language model with zero shot or few shot prompting. The remarkable achievements of PFM have brought significant breakthroughs to various fields of AI. Numerous studies have proposed different methods, raising the demand for an updated survey. This study provides a comprehensive review of recent research advancements, challenges, and opportunities for PFMs in text, image, graph, as well as other data modalities. The review covers the basic components and existing pretraining methods used in natural language processing, computer vision, and graph learning. Additionally, it explores advanced PFMs used for different data modalities and unified PFMs that consider data quality and quantity. The review also discusses research related to the fundamentals of PFMs, such as model efficiency and compression, security, and privacy. Finally, the study provides key implications, future research directions, challenges, and open problems in the field of PFMs. Overall, this survey aims to shed light on the research of the PFMs on scalability, security, logical reasoning ability, cross-domain learning ability, and the user-friendly interactive ability for artificial general intelligence.
研究动机与目标
- 调查跨 NLP、CV 和 GL 的 PFMs 发展,并追踪它们从 BERT 到 ChatGPT 的历史。
- 分析跨模态的基本组成、训练范式和预训练任务。
- 讨论统一 PFMs、模型效率、压缩、安全性和隐私等前沿主题。
- 指出挑战与待解决的问题,以指导未来对 PFMs 的研究。
- 通过附录提供与 PFMs 相关的评估指标和数据集的指南。
提出的方法
- 描述基于 Transformer 的 PFMs 作为主导架构。
- 对学习机制进行分类(有监督、半监督、SSL、RL)及其在预训练中的作用。
- 总结 NLP 的预训练任务(MLM、DAE、RTD、NSP、SOP)以及 CV/GL 的 SSL 策略。
- 概述模型设计选择(自回归、上下文、置换语言模型)以及关键模型(如 GPT、BERT、XLNet、MPNet)。
- 讨论用于使输出符合人类偏好的指令对齐方法(如 RLHF、连锁思维等)。
- 综合讨论统一 PFMs、效率、压缩、安全性与隐私等前沿主题。
实验结果
研究问题
- RQ1推动 NLP、CV 和 GL 的 PFMs 的核心组成和学习机制是什么?
- RQ2跨模态的预训练任务如何演化以提升表征学习与下游性能?
- RQ3自回归、上下文和置换语言模型之间的设计权衡是什么?
- RQ4跨多模态数据的统一 PFMs 的现状如何、哪些挑战尚存?
- RQ5PFMs 在效率、安全性与隐私方面的开放挑战与未来方向是什么?
主要发现
- Transformer 仍然是支撑跨 NLP、CV 和 GL 的可扩展 PFMs 的核心架构。
- 预训练任务和学习机制(SSL、RL 和有监督信号)支撑强健的特征表示和快速微调。
- NLP 预训练任务包括 MLM、DAE、RTD、NSP 和 SOP,扩展如 XLNet 与 MPNet 引入置换/混合目标。
- 新兴趋势包括能够处理文本、图像和音频的统一 PFMs,以 GPT-4 等模型为代表。
- 前沿主题强调模型效率、压缩、安全性与隐私,解决实际部署与治理方面的关注。
- 本综述强调在可扩展性、跨域学习、推理和交互式 AI 能力方面的未来研究方向。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。