Skip to main content
QUICK REVIEW

[Paper Review] A Comprehensive Survey on Cross-modal Retrieval

Kaiye Wang, Qiyue Yin|arXiv (Cornell University)|Jul 21, 2016
Advanced Image and Video Retrieval Techniques4 references224 citations
TL;DR

This survey classifies cross-modal retrieval methods into real-valued and binary representations, reviews unsupervised/pairwise/rank-based/supervised approaches, summarizes datasets and experimental findings, and outlines open problems and future directions.

ABSTRACT

In recent years, cross-modal retrieval has drawn much attention due to the rapid growth of multimodal data. It takes one type of data as the query to retrieve relevant data of another type. For example, a user can use a text to retrieve relevant pictures or videos. Since the query and its retrieved results can be of different modalities, how to measure the content similarity between different modalities of data remains a challenge. Various methods have been proposed to deal with such a problem. In this paper, we first review a number of representative methods for cross-modal retrieval and classify them into two main groups: 1) real-valued representation learning, and 2) binary representation learning. Real-valued representation learning methods aim to learn real-valued common representations for different modalities of data. To speed up the cross-modal retrieval, a number of binary representation learning methods are proposed to map different modalities of data into a common Hamming space. Then, we introduce several multimodal datasets in the community, and show the experimental results on two commonly used multimodal datasets. The comparison reveals the characteristic of different kinds of cross-modal retrieval methods, which is expected to benefit both practical applications and future research. Finally, we discuss open problems and future research directions.

Motivation & Objective

  • Provide a structured overview of cross-modal retrieval research and its motivation.
  • Classify existing methods into real-valued representation learning and binary cross-modal hashing.
  • Summarize datasets, experimental results, and practical implications to guide future work.
  • Discuss open challenges and future research directions in cross-modal retrieval.

Proposed method

  • Classify cross-modal retrieval methods into four categories: unsupervised, pairwise-based, rank-based, and supervised.
  • Differentiate real-valued representation learning from binary (hashing) approaches.
  • Survey subcategories under each main category (e.g., subspace learning, topic models, deep learning, and metric learning).
  • Present representative algorithms and their core ideas, focusing on how they learn common representations across modalities.

Experimental results

Research questions

  • RQ1What are the main categories and subcategories of cross-modal retrieval methods?
  • RQ2How do unsupervised, pairwise-based, rank-based, and supervised approaches differ in learning cross-modal representations?
  • RQ3What datasets are commonly used to evaluate cross-modal retrieval, and what do the experiments show on those datasets?
  • RQ4What open problems and future directions remain in cross-modal retrieval research?

Key findings

  • The paper provides a taxonomy distinguishing real-valued representation learning and binary cross-modal hashing.
  • It analyzes unsupervised, pairwise-based, rank-based, and supervised methods across subspace learning, topic models, and deep learning.
  • It summarizes representative algorithms and discusses their strengths, limitations, and applicability.
  • It introduces multimodal datasets and reports experimental results on two commonly used datasets to illustrate method characteristics.
  • It discusses open problems and opportunities for future research in cross-modal retrieval.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.