Skip to main content
QUICK REVIEW

[논문 리뷰] BlackMarks: Blackbox Multibit Watermarking for Deep Neural Networks

Huili Chen, Bita Darvish Rouhani|arXiv (Cornell University)|2019. 03. 31.
Advanced Steganography and Watermarking Techniques참고 문헌 32인용 수 41
한 줄 요약

BlackMarks는 DNN의 다중 비트 워터마킹을 위한 첫 엔드-투-엔드 블랙박스 프레임워크를 도입하여, 모델 출력에 이진 서명을 미세조정으로 삽입하고, 질의 키를 통한 추출은 낮은 오버헤드와 높은 강건성으로 수행합니다.

ABSTRACT

Deep Neural Networks have created a paradigm shift in our ability to comprehend raw data in various important fields ranging from computer vision and natural language processing to intelligence warfare and healthcare. While DNNs are increasingly deployed either in a white-box setting where the model internal is publicly known, or a black-box setting where only the model outputs are known, a practical concern is protecting the models against Intellectual Property (IP) infringement. We propose BlackMarks, the first end-to-end multi-bit watermarking framework that is applicable in the black-box scenario. BlackMarks takes the pre-trained unmarked model and the owner's binary signature as inputs and outputs the corresponding marked model with a set of watermark keys. To do so, BlackMarks first designs a model-dependent encoding scheme that maps all possible classes in the task to bit '0' and bit '1' by clustering the output activations into two groups. Given the owner's watermark signature (a binary string), a set of key image and label pairs are designed using targeted adversarial attacks. The watermark (WM) is then embedded in the prediction behavior of the target DNN by fine-tuning the model with generated WM key set. To extract the WM, the remote model is queried by the WM key images and the owner's signature is decoded from the corresponding predictions according to the designed encoding scheme. We perform a comprehensive evaluation of BlackMarks's performance on MNIST, CIFAR10, ImageNet datasets and corroborate its effectiveness and robustness. BlackMarks preserves the functionality of the original DNN and incurs negligible WM embedding runtime overhead as low as 2.054%.

연구 동기 및 목표

  • 블랙-박스 환경(MLaas)에서 DNN에 대한 IP 보호를 위한 동기를 부여합니다.
  • 모델 내부에 의존하지 않고 작동하는 확장 가능하고 다중 비트 워터마킹 프레임워크를 개발합니다.
  • 클래스 출력의 비트 인코딩을 모델 의존적으로 설계하고 표적 적대적 공격을 통해 워터마크 키를 생성합니다.
  • 워터마크를 삽입하기 위해 모델을 미세조정하되 정확도는 보존합니다.
  • 낮은 거짓 양성/거짓 음성으로 견고한 추출 및 검증 절차를 제공합니다.

제안 방법

  • 클래스 평균에 대해 K-means로 두 비트-클러스터로 분류하여 DNN 출력 활성화에 소유자 서명을 인코딩합니다(소프트맥스 전).
  • 인코딩 방식과 맞춰진 표적적대적 공격을 사용하여 워터마크 키 이미지와 레이블을 생성합니다.
  • 표준 교차 엔트로피와 WM 특화 손실을 결합한 정규화된 손실로 사전 학습된 모델을 미세조정하여 서명을 삽입합니다.
  • WM 키를 질의하고 인코딩 방식을 적용하여 모델 예측에서 소유자 서명을 디코딩하고 Bit Error Rate(BER)을 산출합니다.
  • 더 큰 초기 키 세트를 선택하고 표식되지 않은 모델 변형 간 교차 교차를 통해 키 선택 시 WM 키의 전이 가능성을 방지합니다.
  • 일회성 삽입 오버헤드가 최소 2.054%에 불과하고 블랙-박스 추출 비용에 대한 효율성 분석을 제공합니다.

실험 결과

연구 질문

  • RQ1블랙-박스 DNN 워터마크가 다중 비트 용량으로 IP 보호를 강화할 수 있습니까?
  • RQ2정확도를 해치지 않으면서 다중 비트 워터마킹을 지지하기 위해 출력물을 비트로 매핑하는 모델 의존적 인코딩은 어떻게 설계할 수 있을까요?
  • RQ3블랙-박스 설정에서 이러한 워터마크가 미세조정, 가지치기 및 재덮어쓰기에 대해 얼마나 강인합니까?
  • RQ4블랙-박스 시나리오에서 워터마크 무결성과 신뢰성을 어떻게 정량화하고 보장할 수 있을까요?

주요 결과

  • BlackMarks는 MNIST, CIFAR-10, ImageNet에서 지정된 키 세트를 사용할 때 삽입 이후 BER가 0인 높은 워터마크 탐지 성능을 달성합니다.
  • Framework는 MNIST, CIFAR-10, ImageNet에 대해 각각 최대 95%, 80%, 90%의 매개변수 가지치기를 허용하며 탐지를 저해하지 않는 반면(과도한 가지치기는 정확도에 해를 끼칩니다).
  • 모델 미세조정(실험에서 최대 100 에폭) 후에도 워터마크가 탐지 가능하며 모든 벤치마크에 대해 BER가 0입니다.
  • 새 워터마크로 덮어쓰더라도 원래 워터마크를 복구하는 것이 불가능하지 않으며 BER은 여전히 0입니다.
  • 삽입은 실행 시간 오버헤드를 낮게 유지하며(최소 2.054%), 워터마킹 방식은 적대적 강건성도 향상시키는 것으로 보이고, 여러 공격 하에서 정확도가 증가하는 부가 이점이 있다고 저자들은 언급합니다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.