[Paper Review] From GPT-4 to Gemini and Beyond: Assessing the Landscape of MLLMs on Generalizability, Trustworthiness and Causality through Four Modalities
This paper evaluates the generalizability, trustworthiness, and causal reasoning capabilities of 230 manually designed cases across 8 multi-modal models—including GPT-4 and Gemini—spanning text, code, image, and video modalities. It identifies 14 empirical findings revealing critical limitations in both proprietary and open-source MLLMs, particularly in causal reasoning and robustness across modalities, offering actionable insights for improving model reliability in real-world applications.
Multi-modal Large Language Models (MLLMs) have shown impressive abilities in generating reasonable responses with respect to multi-modal contents. However, there is still a wide gap between the performance of recent MLLM-based applications and the expectation of the broad public, even though the most powerful OpenAI's GPT-4 and Google's Gemini have been deployed. This paper strives to enhance understanding of the gap through the lens of a qualitative study on the generalizability, trustworthiness, and causal reasoning capabilities of recent proprietary and open-source MLLMs across four modalities: ie, text, code, image, and video, ultimately aiming to improve the transparency of MLLMs. We believe these properties are several representative factors that define the reliability of MLLMs, in supporting various downstream applications. To be specific, we evaluate the closed-source GPT-4 and Gemini and 6 open-source LLMs and MLLMs. Overall we evaluate 230 manually designed cases, where the qualitative results are then summarized into 12 scores (ie, 4 modalities times 3 properties). In total, we uncover 14 empirical findings that are useful to understand the capabilities and limitations of both proprietary and open-source MLLMs, towards more reliable downstream multi-modal applications.
Motivation & Objective
- To investigate the gap between high-performing MLLMs and public expectations despite deployment of leading models like GPT-4 and Gemini.
- To evaluate the generalizability, trustworthiness, and causal reasoning capabilities of recent proprietary and open-source MLLMs across four modalities: text, code, image, and video.
- To enhance transparency and reliability of MLLMs by identifying systematic limitations through qualitative, case-based analysis.
- To provide actionable insights for improving the robustness and trustworthiness of MLLMs in downstream multi-modal applications.
Proposed method
- Designed 230 manually curated, multi-modal cases covering text, code, image, and video inputs to probe model behavior.
- Evaluated 8 models: GPT-4, Gemini (closed-source), and 6 open-source LLMs and MLLMs across 12 scores (4 modalities × 3 properties: generalizability, trustworthiness, causality).
- Employed qualitative analysis to assess model responses on correctness, consistency, and reasoning depth across diverse scenarios.
- Categorized findings into 14 empirical observations based on recurring patterns in model failures and strengths.
- Focused on causal reasoning by designing cases requiring inference beyond surface-level associations.
- Used a standardized scoring framework to compare performance across models and modalities systematically.
Experimental results
Research questions
- RQ1How do proprietary and open-source MLLMs perform in generalizability across text, code, image, and video modalities?
- RQ2To what extent do MLLMs demonstrate trustworthiness in generating consistent, factually accurate, and reliable responses?
- RQ3Can MLLMs reason causally in multi-modal contexts, especially when faced with counterfactual or ambiguous inputs?
- RQ4What are the key failure modes and limitations in MLLM reasoning that persist across both closed- and open-source models?
- RQ5How do model capabilities vary across modalities, and which modalities expose the most significant gaps in reasoning and reliability?
Key findings
- GPT-4 and Gemini show superior performance in generalizability and trustworthiness across text and image modalities, but still exhibit significant failures in causal reasoning tasks.
- Open-source MLLMs consistently underperform compared to GPT-4 and Gemini, especially in complex multi-modal reasoning and code understanding.
- Causal reasoning remains a major weakness across all models, with over 60% of causal reasoning cases resulting in incorrect or speculative answers.
- Model trustworthiness varies widely: some models generate plausible but factually incorrect responses, particularly in video and code modalities.
- Generalizability is most robust in text and image tasks, but degrades significantly in video and code, where models often misinterpret temporal or syntactic structures.
- A recurring failure pattern is the model’s tendency to overfit to superficial cues rather than reasoning about underlying causal relationships, even in high-performing models.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.