[論文レビュー] On Attribution of Deepfakes
本稿では、生成モデルの出力画像とそのランダムシードの間のマッピングを活用して、深フェイクをその元となる生成モデルに確率的に帰属づける、シード再構築に基づく手法を提案する。顔合成タスクにおいて97.62%の帰属付け精度を達成した。この手法は、摂動に対して頑健であり、人間が解釈可能な結果を提示する一方で、モデル開発者を保護するための妥当な否認可能性( plausible deniability )を提供する枠組みも導入している。
Progress in generative modelling, especially generative adversarial networks, have made it possible to efficiently synthesize and alter media at scale. Malicious individuals now rely on these machine-generated media, or deepfakes, to manipulate social discourse. In order to ensure media authenticity, existing research is focused on deepfake detection. Yet, the adversarial nature of frameworks used for generative modeling suggests that progress towards detecting deepfakes will enable more realistic deepfake generation. Therefore, it comes at no surprise that developers of generative models are under the scrutiny of stakeholders dealing with misinformation campaigns. At the same time, generative models have a lot of positive applications. As such, there is a clear need to develop tools that ensure the transparent use of generative modeling, while minimizing the harm caused by malicious applications. Our technique optimizes over the source of entropy of each generative model to probabilistically attribute a deepfake to one of the models. We evaluate our method on the seminal example of face synthesis, demonstrating that our approach achieves 97.62% attribution accuracy, and is less sensitive to perturbations and adversarial examples. We discuss the ethical implications of our work, identify where our technique can be used, and highlight that a more meaningful legislative framework is required for a more transparent and ethical use of generative modeling. Finally, we argue that model developers should be capable of claiming plausible deniability and propose a second framework to do so -- this allows a model developer to produce evidence that they did not produce media that they are being accused of having produced.
研究の動機と目的
- 誤った深フェイクの使用が拡大する誤情報キャンペーンや社会的操作の脅威に対処すること。
- 合成メディアをその生成モデルにまで遡れるフォレンジック技術を開発し、透明性と監査可能性を高めること。
- GAN の学習における敵対的フィードバックループの影響により本質的に脆いとされる深フェイク検出に依存するのを減らすこと。
- 誤って悪意あるコンテンツを生成したと非難された際、モデル開発者が事実を立証できる仕組みを提供すること。
- 生成AIの責任ある使用を規定する統合的技術的・法的・倫理的枠組みを提唱すること。
提案手法
- 各候補となる生成器に対して、与えられた画像を最もよく再現するランダムシードを再構築する最適化問題を定式化する。
- 生成器がエントロピー(シード)から出力画像へと学習したマッピングを活用し、確率的に元モデルを特定する。
- 画像を最もよく再現するシードを選択することで、確率的再構築アプローチを採用し、最も可能性の高い出典を特定する。
- ノイズや敵対的摂動を加えた状況下でも性能を評価し、先行する検出手法と比較することで、頑健性を検証する。
- 付録 D に、モデル出力の記録と内容生成の立証可能な否認を可能にするハイパーレジストリベースのシステムを提案する。
- 帰属付けのためのフォレンジックトレーシングと、モデル開発者を保護するための妥当な否認可能性の両方を提供する二重枠組みを導入する。
実験結果
リサーチクエスチョン
- RQ1シード再構築を用いることで、合成画像をその元となる生成モデルに信頼性高く帰属づけることができるか?
- RQ2この帰属付け手法は、敵対的摂動や生成後の操作に対してどれほど頑健か?
- RQ3人間の専門家は帰属付け結果をどれくらい理解し、合意できるか?
- RQ4モデル開発者は、特定の深フェイクの生成に関与していないことを立証できるか?
- RQ5帰属付けの限界は何か。また、妥当な否認可能性はこれらの限界をどのように緩和できるか?
主な発見
- 提案手法は、顔合成ベンチマークにおいて97.62%の帰属付け精度を達成しており、訓練ステップが僅か数ステップしか差がない GAN 間の区別にも成功している。
- 人間参加者が帰属付け結果に同意した割合は93.7%であり、強い解釈可能性と信頼性を示している。
- 先行する深フェイク検出手法と比較して、敵対的摂動に対して感受性が低く、より優れた頑健性を示している。
- 生成後の操作は帰属付けの信頼性を低下させる可能性があり、モデル分析のみで厳密な整合性を保証できるわけではないことを示している。
- 一部の GAN は、高エントロピー入力を用いて任意の画像を生成可能であり、帰属付けにおける厳密な整合性の実現が不可能である可能性を示している。
- モデル出力のログ記録にハイパーレジストリを活用する仕組みにより、立証可能な妥当な否認が可能となり、法的・倫理的責任の担保が可能になる。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。