Skip to main content
QUICK REVIEW

[論文レビュー] Open-book Video Captioning with Retrieve-Copy-Generate Network

Ziqi Zhang, Zhongang Qi|arXiv (Cornell University)|Mar 9, 2021
Multimodal Machine Learning Applications参考文献 33被引用数 8
ひとこと要約

本稿では、ビデオコーパスから関連する文を検索して生成をガイドする手法として、Open-book Video Captioningという新しいパラダイムを提案する。Retrieve-Copy-Generate (RCG) ネットワークは、クロスモーダル検索器を用いてビデオ関連の文を取得し、コピーメカニズムジェネレータを用いて動的にコピーまたは生成されたキャプションを出力する。MSR-VTTおよびVATEXの両データセットでSOTAを達成し、それぞれ52.9 CIDErおよび57.5 CIDErのスコアを達成した。

ABSTRACT

Due to the rapid emergence of short videos and the requirement for content understanding and creation, the video captioning task has received increasing attention in recent years. In this paper, we convert traditional video captioning task into a new paradigm, \ie, Open-book Video Captioning, which generates natural language under the prompts of video-content-relevant sentences, not limited to the video itself. To address the open-book video captioning problem, we propose a novel Retrieve-Copy-Generate network, where a pluggable video-to-text retriever is constructed to retrieve sentences as hints from the training corpus effectively, and a copy-mechanism generator is introduced to extract expressions from multi-retrieved sentences dynamically. The two modules can be trained end-to-end or separately, which is flexible and extensible. Our framework coordinates the conventional retrieval-based methods with orthodox encoder-decoder methods, which can not only draw on the diverse expressions in the retrieved sentences but also generate natural and accurate content of the video. Extensive experiments on several benchmark datasets show that our proposed approach surpasses the state-of-the-art performance, indicating the effectiveness and promising of the proposed paradigm in the task of video captioning.

研究の動機と目的

  • 従来のビデオキャプションの限界、すなわち視覚的入力に依存するのみで、一般的なキャプションしか生成しないという点を是正すること。
  • エンドツーエンドモデルの固定された知識ドメインを克服し、推論時に外部で動的かつ柔軟にアクセス可能な知識ソースを活用できるようにすること。
  • 生成時に視覚的特徴と取得したテキスト表現を統合することで、キャプションの多様性と正確性を向上させること。
  • リtrievalと生成のコンponentを別々にトレーニング可能とする柔軟性と拡張性に優れたフレームワークの開発

提案手法

  • ビデオからテキストへの検索器は、バイエンコーダー構造に基づき、動きと外観の特徴を併用して大規模なテキストコーパスから意味的に関連する文を検索する。
  • 検索器は、ビデオクリップと関連する自然言語文をマッチングするように学習され、主に行動や物体といったキービジュー概念に焦点を当てる。
  • コピーメカニズムジェネレータは、新しい単語を生成するか、複数の取得済み文から直接表現をコピーフェーズを動的に決定する。
  • ジェネレータは視覚的特徴と取得済み文の両方に注目するため、事実の正確性と言語の多様性を組み合わせたハイブリッド生成が可能になる。
  • RCGフレームワークはエンドツーエンドトレーニングまたは検索器とジェネレータの別個トレーニングをサポートし、柔軟性とスケーラビリティを向上させる。
  • モデルは標準的なビデオ特徴を用いてトレーニングデータでファインチューニングされ、標準的なキャプション損失関数により最適化される。
Figure 1: Pipeline comparison of the existing methods and our method. Our generation is produced based on not only video content but also the cues of multi-retrieved sentences searched from text corpus by a cross-modal retriever. Pluggable retriever provides guidance and expansion for the generating
Figure 1: Pipeline comparison of the existing methods and our method. Our generation is produced based on not only video content but also the cues of multi-retrieved sentences searched from text corpus by a cross-modal retriever. Pluggable retriever provides guidance and expansion for the generating

実験結果

リサーチクエスチョン

  • RQ1外部から取得したテキスト表現が、ビデオに表示されていない内容を越えて、生成キャプションの質と多様性を向上させることができるか?
  • RQ2大規模なテキストコーパスから、ビデオキャプションに適した意味的に関連する文をクロスモーダル検索器がどれほど効果的に特定できるか?
  • RQ3複数の取得済み文から動的にコピーフェーズを採用することで、純粋な生成に比べてキャプションの正確性と流暢さがどの程度向上するか?
  • RQ4RCGフレームワークはエンドツーエンドまたは別個トレーニングで学習可能か?また、そのトレーニング方法が性能に与える影響は?
  • RQ5オープンブックパラダイムは、特に意味的重複性の高い動画を含む多様な動画データセットに一般化しやすいか?

主な発見

  • RCGモデルはMSR-VTTで52.9 CIDErのスコアを達成し、以前のSOTAであるORG-TRLを3.9%相対的に上回った。
  • VATEXでは57.5 CIDErのスコアを記録し、ORG-TRLに比べ15.7%の相対的改善を示し、意味的重複性の高い動画でも優れた性能を発揮した。
  • +FixRetバージョンは、MSR-VTTで52.3 CIDEr、VATEXで56.8 CIDErを達成し、固定された事前学習済み検索器の有効性を示した。
  • +TrainRetバージョンは、検索器とジェネレータを共同でファインチューニングすることで、最高の性能を達成し、エンドツーエンド最適化の利点を確認した。
  • 定性的な分析により、モデルが「マスクから息を吐いている」のような有用なフレーズを効果的にコピーフェーズで再利用し、一般的なキャプション「フリスビーのゲームをやっている」を「キャッチをやっている」に修正していることが確認された。
  • ヒートマップの可視化により、コピーメカニズムが関連するキーワードに集中していることが示され、検索の注目重みは「スキューバダイバー」や「海の水」のような顕著な概念に集中していた。
Figure 2: Overview of the proposed Retrieve-Copy-Generate Network for Open-book Video Captioning. The left side is the pipeline of our method, which consists of two components: the Video-to-Text Retriever that searches for the video-content-relevant sentences from the corpus containing all the sente
Figure 2: Overview of the proposed Retrieve-Copy-Generate Network for Open-book Video Captioning. The left side is the pipeline of our method, which consists of two components: the Video-to-Text Retriever that searches for the video-content-relevant sentences from the corpus containing all the sente

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。