Skip to main content
QUICK REVIEW

[論文レビュー] Deep Semantic Multimodal Hashing Network for Scalable Multimedia Retrieval.

Lu Jin, Jinhui Tang|arXiv (Cornell University)|Jan 9, 2019
Advanced Image and Video Retrieval Techniques参考文献 75被引用数 5
ひとこと要約

本稿では、相互モダリティ類似性、内部モダリティの意味的ラベル、およびビットバランス制約を保持することで、モダリティ固有のハッシュ関数を統合的に学習する、統合的ディープラーニングフレームワークであるDeep Semantic Multimodal Hashing Network (DSMHN) を提案する。訓練中に意味的ラベルをハッシュコードに直接埋め込むことで、DSMHNは3つのベンチマークデータセットにおいて最先端の手法を上回る優れたリtrieval性能を達成する。

ABSTRACT

Hashing has been widely applied to multimodal retrieval on large-scale multimedia data due to its efficiency in computation and storage. Particularly, deep hashing has received unprecedented research attention in recent years, owing to its perfect retrieval performance. However, most of existing deep hashing methods learn binary hash codes by preserving the similarity relationship while without exploiting the semantic labels of data points, which result in suboptimal binary codes. In this work, we propose a novel Deep Semantic Multimodal Hashing Network for scalable multimodal retrieval. In DSMHN, two sets of modality-specific hash functions are jointly learned by explicitly preserving both the inter-modality similarities and the intra-modality semantic labels. Specifically, with the assumption that the learned hash codes should be optimal for task-specific classification, two stream networks are jointly trained to learn the hash functions by embedding the semantic labels on the resultant hash codes. Different from previous deep hashing methods, which are tied to some particular forms of loss functions, the proposed deep hashing framework can be flexibly integrated with different types of loss functions. In addition, the bit balance property is investigated to generate binary codes with each bit having 50% probability to be 1 or -1. Moreover, a unified deep multimodal hashing framework is proposed to learn compact and high-quality hash codes by exploiting the feature representation learning, inter-modality similarity preserving learning, semantic label preserving learning and hash functions learning with bit balanced constraint simultaneously. We conduct extensive experiments for both unimodal and cross-modal retrieval tasks on three widely-used multimodal retrieval datasets. The experimental result demonstrates that DSMHN significantly outperforms state-of-the-art methods.

研究の動機と目的

  • 既存のディープハッシング手法がバイナリコード学習中に意味的ラベルを無視するという制限に対処すること。
  • ハッシュコード学習プロセスに意味的ラベル情報を明示的に組み込むことで、リtrieval精度を向上させること。
  • 固定アーキテクチャにとどまらず、さまざまな損失関数と互換性を持つ柔軟なディープハッシングフレームワークを開発すること。
  • 各ビットが1または-1である確率が約50%になるように、ハッシュコードのバランスを保証することで、コード品質を向上させること。
  • 特徴表現学習、類似性保持、意味的ラベル学習、およびビットバランスを1つのエンドツーエンドフレームワークに統合すること。

提案手法

  • 共有およびモダリティ固有の特徴を用いて、2本のストリームネットワークを同時に訓練し、モダリティ固有のハッシュ関数を学習する。
  • 訓練中に意味的ラベルを直接ハッシュコードに埋め込み、意味的に意味のあるバイナリコードの学習を導く。
  • 同じ意味的類似性を持つクロスモダリティペアのハッシュコード間の距離を最小化することで、相互モダリティ類似性を保持する。
  • 内部モダリティの意味的ラベルの保持を強制するために、ハッシュコードをその対応する意味的埋め込みと一致させる。
  • 各ビットが1または-1をとる確率がほぼ等しくなるように、ビットバランス制約を適用する。
  • フレームワークは柔軟に設計されており、さまざまな損失関数と互換性があり、多様な最適化目的との統合を可能にする。

実験結果

リサーチクエスチョン

  • RQ1意味的ラベルをディープハッシュコード学習に明示的に組み込むことで、マルチモーダルリtrievalタスクにおけるリtrieval性能が向上するか?
  • RQ2相互モダリティ類似性と内部モダリティ意味的ラベルの共同学習が、学習されたバイナリコードの品質にどのように影響するか?
  • RQ3ビットバランスを強制することで、ディープマルチモーダルハッシングモデルの性能がどの程度向上するか?
  • RQ4提案されたフレームワークは、性能を損なわせることなく、さまざまな損失関数に柔軟に拡張可能か?
  • RQ5DSMHNは、ユニモーダルおよびクロスモーダルリtrievalの両状況において、最先端の手法と比較してどのように差をつけるか?

主な発見

  • DSMHNは、3つの広く使われているマルチモーダルリtrieバルデータセットにおいて、既存の最先端のディープハッシング手法を著しく上回る。
  • ハッシュコード学習への意味的ラベルの統合により、より判別力があり意味的に意味のあるバイナリコードが得られる。
  • ビットバランス制約は、一般化性能の向上とハッシュコードのより均一な分布に寄与する。
  • 特徴学習、類似性保持、意味的ラベル整合の共同最適化により、ユニモーダルおよびクロスモーダルタスクの両方でリtrieval精度が向上する。
  • 提案されたフレームワークは、性能の低下を伴わずに、さまざまな損失関数との統合を可能にする強力な一般化能力と柔軟性を示している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。