[论文解读] WenLan 2.0: Make AI Imagine via a Multimodal Foundation Model.
WenLan 2.0 引入了一种多模态基础模型,该模型通过在弱相关联的网络爬取视觉与文本数据上进行自监督预训练,实现了在多样化认知任务上的强大零样本泛化能力。它展现出先进的想象力和常识推理能力,标志着向通用人工智能迈出了重要一步。
The fundamental goal of artificial intelligence (AI) is to mimic the core cognitive activities of human including perception, memory, and reasoning. Although tremendous success has been achieved in various AI research fields (e.g., computer vision and natural language processing), the majority of existing works only focus on acquiring single cognitive ability (e.g., image classification, reading comprehension, or visual commonsense reasoning). To overcome this limitation and take a solid step to artificial general intelligence (AGI), we develop a novel foundation model pre-trained with huge multimodal (visual and textual) data, which is able to be quickly adapted for a broad class of downstream cognitive tasks. Such a model is fundamentally different from the multimodal foundation models recently proposed in the literature that typically make strong semantic correlation assumption and expect exact alignment between image and text modalities in their pre-training data, which is often hard to satisfy in practice thus limiting their generalization abilities. To resolve this issue, we propose to pre-train our foundation model by self-supervised learning with weak semantic correlation data crawled from the Internet and show that state-of-the-art results can be obtained on a wide range of downstream tasks (both single-modal and cross-modal). Particularly, with novel model-interpretability tools developed in this work, we demonstrate that strong imagination ability (even with hints of commonsense) is now possessed by our foundation model. We believe our work makes a transformative stride towards AGI and will have broad impact on various AI+ fields (e.g., neuroscience and healthcare).
研究动机与目标
- 为克服现有多模态基础模型依赖预训练数据中强语义对齐的局限性。
- 开发一种能够在多样化单模态与跨模态认知任务中实现泛化的基础模型。
- 在预训练阶段无需精确图像-文本对齐的情况下,实现人工智能模型的强大想象力与常识推理能力。
- 证明在弱相关互联网数据上进行自监督学习,对于构建可泛化的多模态表征的有效性。
- 提供新颖的模型可解释性工具,以分析和验证模型的推理与想象力能力。
提出的方法
- 使用从互联网上爬取的大规模弱相关联多模态数据,通过自监督学习预训练基础模型。
- 利用图像与文本之间的弱语义关联,而非要求精确对齐,从而提升模型的鲁棒性与泛化能力。
- 设计一种多模态架构,联合编码视觉与文本输入,以支持多样化的下游任务。
- 引入新颖的模型可解释性工具,用于分析内部表征与推理过程,尤其针对想象力与常识推理。
- 在广泛的下游任务上微调预训练模型,包括图像分类、视觉问答与文本生成。
- 利用对比学习与掩码建模目标,在预训练阶段无需成对标注即可学习联合表征。
实验结果
研究问题
- RQ1在弱相关网络数据上预训练的基础模型,是否能在多样化认知任务中实现优异性能?
- RQ2此类模型在无显式监督的情况下,其想象力与常识推理能力能达到何种程度?
- RQ3与监督学习或强对齐预训练相比,基于弱对齐数据的自监督学习在泛化能力方面表现如何?
- RQ4新颖的可解释性工具是否能有效揭示并验证模型的内部推理与想象力能力?
- RQ5在弱相关数据上进行预训练,对下游零样本与少样本迁移性能有何影响?
主要发现
- WenLan 2.0 模型在广泛的下游认知任务中实现了最先进性能,涵盖单模态与跨模态基准测试。
- 该模型展现出强大的零样本泛化能力,可在无需微调的情况下有效泛化至未见任务。
- 通过新颖的可解释性工具,模型表现出想象力与常识推理的证据,即使在提示信息极少的情况下亦然。
- 在弱相关数据上进行自监督预训练,可生成稳健且可泛化的多模态表征,其性能优于依赖强对齐预训练数据的模型。
- 尽管预训练阶段未使用精确图像-文本对齐,该模型在标准基准测试上的表现仍与现有模型相当或更优。
- 可解释性工具揭示,该模型能够生成合理且上下文相关的视觉与文本推理,表明其具备涌现的认知类行为。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。