Skip to main content
QUICK REVIEW

[论文解读] Deep Learning Techniques for Future Intelligent Cross-Media Retrieval

Sadaqat Ur Rehman, Muhammad Waqas|arXiv (Cornell University)|Jul 21, 2020
Advanced Image and Video Retrieval Techniques参考文献 182被引用 190
一句话总结

本论文提供了对跨媒体检索中深度学习方法的全面综述,提出基于表示、对齐和翻译的分类法,并回顾数据集和挑战。

ABSTRACT

With the advancement in technology and the expansion of broadcasting, cross-media retrieval has gained much attention. It plays a significant role in big data applications and consists in searching and finding data from different types of media. In this paper, we provide a novel taxonomy according to the challenges faced by multi-modal deep learning approaches in solving cross-media retrieval, namely: representation, alignment, and translation. These challenges are evaluated on deep learning (DL) based methods, which are categorized into four main groups: 1) unsupervised methods, 2) supervised methods, 3) pairwise based methods, and 4) rank based methods. Then, we present some well-known cross-media datasets used for retrieval, considering the importance of these datasets in the context in of deep learning based cross-media retrieval approaches. Moreover, we also present an extensive review of the state-of-the-art problems and its corresponding solutions for encouraging deep learning in cross-media retrieval. The fundamental objective of this work is to exploit Deep Neural Networks (DNNs) for bridging the "media gap", and provide researchers and developers with a better understanding of the underlying problems and the potential solutions of deep learning assisted cross-media retrieval. To the best of our knowledge, this is the first comprehensive survey to address cross-media retrieval under deep learning methods.

研究动机与目标

  • 提出一个以表示、对齐和翻译为核心的跨媒体检索挑战分类法。
  • 审视跨媒体检索中的深度学习方法,覆盖无监督、监督、成对和基于排序的范式。
  • 评述知名的跨媒体数据集及其对基于DL的检索方法的适用性。
  • 识别跨媒体基于DL的检索中的当前问题、空白与未来研究机会。

提出的方法

  • 定义跨媒体检索挑战的分类:表示、对齐和翻译。
  • 将基于DL的跨媒体检索方法分为四类:无监督、监督、成对和基于排序。
  • 调研跨媒体数据集并总结其特征及与DL方法的相关性。
  • 讨论前沿问题及为弥合媒体差距提出的基于DL的解决方案。
  • 主张端到端DL框架和多模态表示作为实现跨媒体检索的关键因素。

实验结果

研究问题

  • RQ1在使用深度学习时,跨媒体检索的关键挑战(表示、对齐、翻译)是什么?
  • RQ2基于DL的方法(无监督、监督、成对、基于排序)如何应对这些挑战?
  • RQ3哪些数据集最能支持对基于DL的跨媒体检索方法的评估和推进?
  • RQ4基于DL的跨媒体检索的主要差距和未来方向是什么?

主要发现

  • 提出一个新颖的分类法,涵盖DL基于跨媒体检索中的表示、对齐和翻译。
  • 提供了无监督、监督、成对和基于排序方法的最新DL综述。
  • 对广泛使用的跨媒体数据集及其在DL评估中的优缺点进行了详细评述。
  • 强调最前沿的问题与机遇,为跨媒体DL检索的未来研究提供方向。
  • 强调端到端的DL模型和多模态表示是弥合媒体差距的关键。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。