Skip to main content
QUICK REVIEW

[Paper Review] Integration of Mixture of Experts and Multimodal Generative AI in Internet of Vehicles: A Survey

Minrui Xu, Dusit Niyato|arXiv (Cornell University)|Apr 25, 2024
Big Data Technologies and ApplicationsDecision Sciences3 citations
TL;DR

This survey proposes integrating Mixture of Experts (MoE) and multimodal Generative AI (GAI) to enable Artificial General Intelligence (AGI) in the Internet of Vehicles (IoV), leveraging distributed, collaborative AI for autonomous decision-making. The approach enhances cognitive functions through parameter-efficient, privacy-preserving, and low-carbon AI models that support real-time perception, simulation, and planning across connected vehicles.

ABSTRACT

Generative AI (GAI) can enhance the cognitive, reasoning, and planning capabilities of intelligent modules in the Internet of Vehicles (IoV) by synthesizing augmented datasets, completing sensor data, and making sequential decisions. In addition, the mixture of experts (MoE) can enable the distributed and collaborative execution of AI models without performance degradation between connected vehicles. In this survey, we explore the integration of MoE and GAI to enable Artificial General Intelligence in IoV, which can enable the realization of full autonomy for IoV with minimal human supervision and applicability in a wide range of mobility scenarios, including environment monitoring, traffic management, and autonomous driving. In particular, we present the fundamentals of GAI, MoE, and their interplay applications in IoV. Furthermore, we discuss the potential integration of MoE and GAI in IoV, including distributed perception and monitoring, collaborative decision-making and planning, and generative modeling and simulation. Finally, we present several potential research directions for facilitating the integration.

Motivation & Objective

  • To enable full autonomy in IoV through the fusion of Mixture of Experts (MoE) and multimodal Generative AI (GAI) for enhanced cognitive and reasoning capabilities.
  • To address resource constraints and high mobility in IoV by enabling distributed, collaborative AI execution via MoE architecture.
  • To support advanced IoV applications such as autonomous driving, traffic management, and environment monitoring through generative modeling and data augmentation.
  • To identify critical research challenges in privacy, energy efficiency, and model adaptability for scalable AGI deployment in IoV.
  • To propose future research directions for communication-efficient, parameter-efficient, and retrieval-augmented MoE-GAI systems in dynamic vehicular networks.

Proposed method

  • Utilizes Mixture of Experts (MoE) to partition AI model parameters into specialized expert networks, enabling efficient, distributed inference across vehicles and RSUs.
  • Employs multimodal GAI models—such as diffusion models, VAEs, GANs, and transformers—to process and generate diverse data modalities including images, text, and point clouds.
  • Integrates temporal modeling with key-frame controllers and sliding window techniques to maintain cross-frame consistency in video generation for autonomous driving.
  • Applies retrieval-augmented generation (RAG) to enhance GAI output relevance by retrieving contextually accurate data during generation, improving decision accuracy.
  • Leverages privacy-preserving techniques like federated learning, homomorphic encryption, and secure multi-party computation for collaborative inference without exposing raw data.
  • Proposes low-carbon multimodal perception by designing energy-efficient AI models that minimize computational energy use while maintaining performance.

Experimental results

Research questions

  • RQ1How can MoE and multimodal GAI be jointly designed to enable scalable, distributed, and collaborative AI in high-mobility IoV environments?
  • RQ2What are the key challenges in ensuring privacy and data security when performing collaborative inference across interconnected vehicles?
  • RQ3How can energy-efficient AI models be developed to support sustainable, low-carbon multimodal perception in IoV systems?
  • RQ4In what ways can parameter-efficient fine-tuning enhance the adaptability of local experts in dynamic IoV scenarios without retraining the full model?
  • RQ5How can retrieval-augmented generation improve the contextual accuracy and real-time relevance of GAI outputs in IoV applications like traffic prediction and autonomous navigation?

Key findings

  • The integration of MoE and multimodal GAI enables scalable, distributed AI execution in IoV, supporting real-time cognitive functions without performance degradation under high mobility.
  • Privacy-preserving collaborative inference using federated learning and homomorphic encryption allows secure model updates across vehicles while protecting raw data.
  • Low-carbon multimodal perception models can significantly reduce energy consumption in IoV deployments, supporting sustainability goals and energy-efficient operations.
  • Parameter-efficient fine-tuning enables rapid adaptation of pre-trained local experts to new IoV scenarios with minimal computational overhead.
  • Retrieval-augmented generation enhances the contextual accuracy of GAI outputs by incorporating real-time, relevant data during generation, improving reliability in dynamic environments.
  • The convergence of MoE and GAI presents a transformative pathway toward AGI in IoV, enabling autonomous, intelligent, and adaptive transportation systems with minimal human supervision.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.