[Paper Review] Visual Knowledge in the Big Model Era: Retrospect and Prospect
This paper reviews the evolution of visual knowledge as a unified, interpretable representation of visual concepts, relations, operations, and reasoning, proposing that integration with large foundation models can overcome key limitations in transparency, generalization, and reasoning. It advocates for leveraging big models to extract, enrich, and complement visual knowledge, enabling more trustworthy and capable next-generation AI systems.
Visual knowledge is a new form of knowledge representation that can encapsulate visual concepts and their relations in a succinct, comprehensive, and interpretable manner, with a deep root in cognitive psychology. As the knowledge about the visual world has been identified as an indispensable component of human cognition and intelligence, visual knowledge is poised to have a pivotal role in establishing machine intelligence. With the recent advance of Artificial Intelligence (AI) techniques, large AI models (or foundation models) have emerged as a potent tool capable of extracting versatile patterns from broad data as implicit knowledge, and abstracting them into an outrageous amount of numeric parameters. To pave the way for creating visual knowledge empowered AI machines in this coming wave, we present a timely review that investigates the origins and development of visual knowledge in the pre-big model era, and accentuates the opportunities and unique role of visual knowledge in the big model era.
Motivation & Objective
- To analyze the theoretical and methodological foundations of visual knowledge from cognitive psychology and AI research.
- To identify the limitations of current large AI models, including opacity, hallucination, and high resource demands.
- To explore how visual knowledge can mitigate weaknesses in big models by enabling structured, interpretable reasoning about visual concepts.
- To propose a synergistic framework where big models serve as sources, tools, and complements for visual knowledge creation and enhancement.
- To chart future research directions for integrating visual knowledge with foundation models to advance trustworthy and generalizable AI.
Proposed method
- Systematically categorizes visual knowledge along four dimensions: visual concepts, visual relations, visual operations, and visual reasoning.
- Draws on cognitive psychology to ground visual knowledge in human mental imagery and perception, emphasizing interpretability and abstraction.
- Proposes using foundation models (e.g., GPT, SAM) to learn robust visual concepts and relations from massive data without fine-tuning.
- Introduces the use of large language models as external knowledge sources to enrich visual knowledge with semantic, commonsense, and contextual relationships.
- Advocates for advanced knowledge analysis, extraction, and enhancement techniques to localize, validate, and refine knowledge encoded in big models.
- Emphasizes multimodal synergy—leveraging textual and visual modalities jointly to create more holistic, context-aware representations.
Experimental results
Research questions
- RQ1How can visual knowledge serve as a bridge between symbolic reasoning and sub-symbolic deep learning in AI?
- RQ2What are the core components of visual knowledge, and how do they support structured reasoning about visual entities?
- RQ3In what ways can foundation models enhance the acquisition, representation, and refinement of visual knowledge?
- RQ4How can visual knowledge mitigate the opacity, hallucination, and bias issues inherent in large AI models?
- RQ5What are the most promising research directions for integrating visual knowledge with big models to enable trustworthy, generalizable AI?
Key findings
- Visual knowledge provides a unified, interpretable framework for representing visual concepts, their attributes (e.g., shape, motion), and their transformations, enabling structured reasoning.
- Foundation models like SAM and GPT-3 demonstrate the potential to learn visual patterns and world knowledge at scale, serving as powerful engines for visual knowledge extraction.
- Large language models contain significant implicit world knowledge—including semantic and commonsense relationships—that can enrich visual knowledge but remain hidden within model parameters.
- Knowledge extraction from big models remains challenging due to entanglement with noise, bias, and errors, necessitating advanced analysis and refinement techniques.
- Complementary use of textual and visual modalities enhances understanding beyond what either modality can achieve alone, especially for abstract or contextual knowledge (e.g., historical significance of landmarks).
- Integrating visual knowledge with foundation models offers a viable path toward next-generation AI that balances performance, interpretability, and reliability.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.