[논문 리뷰] AdvParams: An Active DNN Intellectual Property Protection Technique via Adversarial Perturbation Based Parameter Encryption
이 논문은 적대적 편향을 사용하여 최소한의 핵심 모델 파라미터를 암호화함으로써 비인가 사용자에게 모델 정확도를 극적으로 낮추면서도, 비밀 키를 가진 정당한 사용자에게는 높은 정확도를 유지하는 활성 딥 뉴럴 네트워크(DNN) 지적재산권 보호 방법인 AdvParams를 제안한다. 이 방법은 재학습이나 하드웨어 지원 없이도 GTSRB에서 최대 87.91%의 정확도 저하를 달성한다.
A well-trained DNN model can be regarded as an intellectual property (IP) of the model owner. To date, many DNN IP protection methods have been proposed, but most of them are watermarking based verification methods where model owners can only verify their ownership passively after the copyright of DNN models has been infringed. In this paper, we propose an effective framework to actively protect the DNN IP from infringement. Specifically, we encrypt the DNN model's parameters by perturbing them with well-crafted adversarial perturbations. With the encrypted parameters, the accuracy of the DNN model drops significantly, which can prevent malicious infringers from using the model. After the encryption, the positions of encrypted parameters and the values of the added adversarial perturbations form a secret key. Authorized user can use the secret key to decrypt the model. Compared with the watermarking methods which only passively verify the ownership after the infringement occurs, the proposed method can prevent infringement in advance. Moreover, compared with most of the existing active DNN IP protection methods, the proposed method does not require additional training process of the model, which introduces low computational overhead. Experimental results show that, after the encryption, the test accuracy of the model drops by 80.65%, 81.16%, and 87.91% on Fashion-MNIST, CIFAR-10, and GTSRB, respectively. Moreover, the proposed method only needs to encrypt an extremely low number of parameters, and the proportion of the encrypted parameters of all the model's parameters is as low as 0.000205%. The experimental results also indicate that, the proposed method is robust against model fine-tuning attack and model pruning attack. Moreover, for the adaptive attack where attackers know the detailed steps of the proposed method, the proposed method is also demonstrated to be robust.
연구 동기 및 목표
- 침해 발생 후에만 소유권을 검증할 수 있는 수동 워터마킹 기법의 한계를 해결하기 위해.
- 침해 발생 이전에 비정상적 사용을 방지할 수 있는 활성 DNN IP 보호 메커니즘을 개발하기 위해.
- 보호 과정에서 재학습이나 하드웨어 의존도를 피하여 계산 오버헤드를 최소화하기 위해.
- 일반적인 공격, 예를 들어 미세조정, 프루닝 및 적응형 공격에 대해 강건성을 확보하기 위해.
제안 방법
- 손실 함수에 대한 가중치에 대한 기울기를 계산하여 모델에서 가장 영향력 있는 파라미터를 식별한다.
- 선택된 파라미터에 대상 적대적 편향을 적용하여 성능 저하를 최대화한다.
- 암호화된 파라미터의 위치와 편향 값이 포함된 비밀 키를 생성한다.
- 정당한 사용자에게는 비밀 키를 사용하여 복호화 과정에서 원래 모델 동작을 복원한다.
- 암호화된 파라미터 수를 최소화함—최대 경우 23개의 파라미터만 암호화—동시에 탐지할 수 없는 편향 크기를 유지한다.
- 모델의 미세조정, 프루닝 및 적응형 공격에 대해 저항력을 입증한다. 이는 공격자가 메커니즘의 작동 원리를 알고 있는 상황에서도 적용된다.
실험 결과
연구 질문
- RQ1적대적 편향을 사용하여 일부 핵심 파라미터만 선택적으로 암호화하여 비인가 사용자에게 DNN 모델을 비활성화시킬 수 있는가?
- RQ2제안된 방법이 비인가 사용자에게는 정확도를 크게 낮추면서도 정당한 사용자에게는 높은 정확도를 유지하는 데 얼마나 효과적인가?
- RQ3모델의 미세조정 및 프루닝 공격에 대해 이 방법이 여전히 강건한가? 이러한 공격은 모델의 功能 복구를 尝試한다.
- RQ4재학습이나 하드웨어 의존도 없이 적용 가능하여 실세계 구현에서 낮은 계산 오버헤드를 유도할 수 있는가?
- RQ5공격자가 암호화 메커니즘의 내부 단계를 모두 알고 있는 적응형 공격 상황에서 이 방법은 어떻게 작동하는가?
주요 결과
- 비인가 사용자에게는 Fashion-MNIST, CIFAR-10, GTSRB에서 각각 80.65%, 81.16%, 87.91%의 테스트 정확도 저하가 발생했다.
- 최대 경우 23개의 파라미터만 암호화되어 총 모델 파라미터의 0.000205%에 불과했다.
- 세 데이터셋에서 정당한 사용자에게는 각각 91.01%, 92.02%, 94.85%의 정확도를 유지했다.
- 모델의 미세조정 공격 후 정확도는 최소 61.66%로 떨어졌고, 프루닝 후에는 17.64%로 감소했다.
- 적응형 공격 상황에서도 공격자가 메커니즘의 내부 단계를 모두 알고 있음에도 불구하고, 정확도는 낮게 유지되었으며, Fashion-MNIST에서는 56.28%, CIFAR-10에서는 62.03%, GTSRB에서는 63.92%였다.
- 재학습이나 하드웨어 지원이 전혀 필요 없어 상용 배포에 실용적이다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.