Skip to main content
QUICK REVIEW

[論文レビュー] On the Adversarial Robustness of Multi-Modal Foundation Models

Christian Schlarmann, Matthias Hein|arXiv (Cornell University)|Aug 21, 2023
Multimodal Machine Learning Applications被引用数 5
ひとこと要約

この論文は、OpenFlamingoのようなマルチモーダル基盤モデルが、視覚入力に対して人間には感知できない悪意ある攻撃に対して極めて脆弱であることを示している。微小な摂動($varepsilon_{\infty} = \nicefrac{1}{255}$)によって、モデルの出力を標的の誤ったキャプションや応答に操作することができる。本研究では、検出されないまま誤情報の拡散やユーザーの操作が可能になることが明らかになった。これにより、展開済みのモデルにおける耐性向上の緊急の必要性が浮き彫りになった。

ABSTRACT

Multi-modal foundation models combining vision and language models such as Flamingo or GPT-4 have recently gained enormous interest. Alignment of foundation models is used to prevent models from providing toxic or harmful output. While malicious users have successfully tried to jailbreak foundation models, an equally important question is if honest users could be harmed by malicious third-party content. In this paper we show that imperceivable attacks on images in order to change the caption output of a multi-modal foundation model can be used by malicious content providers to harm honest users e.g. by guiding them to malicious websites or broadcast fake information. This indicates that countermeasures to adversarial attacks should be used by any deployed multi-modal foundation model.

研究の動機と目的

  • マルチモーダル基盤モデルが、誠実なユーザーが関与する実世界の展開環境において、悪意ある視覚的攻撃に対してどれほど感受性を示すかを調査すること。
  • OpenFlamingoモデルを用いて、画像キャプション生成およびVQAタスクにおける標的攻撃と非標的攻撃の有効性を評価すること。
  • このような脆弱性の実世界への影響を実証すること。具体的には、フェイクニュースの拡散やユーザーを悪意あるコンテンツへ誘導する可能性を含む。
  • 悪意ある第三者による悪用を防ぐために、マルチモーダルモデルにおける耐性メカニズムの必要性を強調すること。
  • 現実的な脅威モデル下でのマルチモーダルモデルにおける悪意ある耐性の評価フレームワークを提供すること。

提案手法

  • 標的および非標的攻撃のための悪意ある摂動を生成するために、$\varepsilon_{\infty} = \nicefrac{1}{255}$ で5000イテレーションを用いたAuto-PGD(APGD)攻撃アルゴリズムを採用した。
  • 攻撃は、ゼロショットのOpenFlamingoモデルを用いて、画像キャプション生成および視覚的質問応答(VQA)タスクの両方で評価された。
  • 性能は、キャプション生成ではCIDErスコア、標的攻撃では攻撃成功確率という標準指標を用いて測定された。
  • 摂動の耐性を評価するために、最小の絶対値を持つ成分から順にゼロに近づけることで、最小限の有効摂動を特定した。
  • 脅威モデルでは、ユーザーによる直接の改ざんではなく、モデルが処理する前に悪意ある第三者が入力画像を操作すると想定した。
  • 実験はCOCOデータセットを用い、ユーザーを操作する能力をテストするために、'Please reset your password'のような標的キャプションが使用された。
Figure 1 : Generated captions on original (left) and adversarially perturbed images (right). We perform a targeted attack on the caption output with $\varepsilon=\nicefrac{{1}}{{255}}$ on the zero-shot model. This could be used for guiding users to a malicious website (top) or fake information (bott
Figure 1 : Generated captions on original (left) and adversarially perturbed images (right). We perform a targeted attack on the caption output with $\varepsilon=\nicefrac{{1}}{{255}}$ on the zero-shot model. This could be used for guiding users to a malicious website (top) or fake information (bott

実験結果

リサーチクエスチョン

  • RQ1人間には感知できない悪意ある摂動は、OpenFlamingoのようなマルチモーダル基盤モデルの出力を操作できるか?
  • RQ2微小な画像変更によって、標的攻撃はどれほどモデルのテキスト出力を制御できるか?
  • RQ3摂動の大きさの一部(例:60%)しか保持しなかった場合、悪意ある摂動はどれほど有効か?
  • RQ4ニュース生成やユーザー誘導といった応用分野において、このような脆弱性がもたらす実世界の影響は何か?
  • RQ5特に$varepsilon_{\infty}$値が小さい場合、攻撃の耐性は異なる脅威モデルにおいてどのように変化するか?

主な発見

  • $varepsilon_{\infty} = \nicefrac{1}{255}$ の悪意ある攻撃は、OpenFlamingoモデルの出力を高く成功確率で操作し、'Please reset your password'のような標的キャプションを生成できた。
  • 摂動の大きさの60%しか保持しなくても、攻撃の成功確率は依然として高く維持された。これは、極めて小さな摂動でも依然として非常に効果的であることを示している。
  • 非標的攻撃ではCIDErスコアが著しく低下した。これは、悪意ある入力下でモデルの性能が著しく低下することを確認している。
  • 摂動は人間には感知できないため、ユーザーの検出なしにモデル出力をすばやく操作するのに理想的である。
  • 本研究では、悪意ある第三者がこれらの脆弱性を悪用し、誤った情報を注入したり、モデル生成テキストを通じてユーザーを有害なコンテンツへ誘導したりする可能性があることが確認された。
  • 結果として、現在のマルチモーダル基盤モデルは、極めて小さな摂動半径であっても、悪意ある入力に対して十分な耐性を備えていないことが示された。
Figure 2 : Generated captions on original and adversarially perturbed images. The perturbations are obtained with a targeted attack using radius $\varepsilon_{q}=\nicefrac{{1}}{{255}}$ and 5000 APGD iterations on the zero-shot model. We show only the original images as the perturbations at this radi
Figure 2 : Generated captions on original and adversarially perturbed images. The perturbations are obtained with a targeted attack using radius $\varepsilon_{q}=\nicefrac{{1}}{{255}}$ and 5000 APGD iterations on the zero-shot model. We show only the original images as the perturbations at this radi

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。