Skip to main content
QUICK REVIEW

[논문 리뷰] Efficient Batch Homomorphic Encryption for Vertically Federated XGBoost

Wuxing Xu, Hao Fan|arXiv (Cornell University)|2021. 12. 08.
Privacy-Preserving Technologies in Data인용 수 10
한 줄 요약

이 논문은 수직 분산 XGBoost를 위한 효율적인 배치 동형 암호화 방법을 제안하며, 일阶 및 이阶 미분을 하나의 값으로 인코딩하여 동형 암호화 및 암호문 전송 비용을 약 50% 감소시킵니다. 이 방법은 SecureBoost 대비 최대 50% 빠른 트리 구축과 30% 이상의 총 런타임 감소를 달성하며, AUC 및 KS 점수로 측정한 정확도 손실은 극히 미미합니다.

ABSTRACT

More and more orgainizations and institutions make efforts on using external data to improve the performance of AI services. To address the data privacy and security concerns, federated learning has attracted increasing attention from both academia and industry to securely construct AI models across multiple isolated data providers. In this paper, we studied the efficiency problem of adapting widely used XGBoost model in real-world applications to vertical federated learning setting. State-of-the-art vertical federated XGBoost frameworks requires large number of encryption operations and ciphertext transmissions, which makes the model training much less efficient than training XGBoost models locally. To bridge this gap, we proposed a novel batch homomorphic encryption method to cut the cost of encryption-related computation and transmission in nearly half. This is achieved by encoding the first-order derivative and the second-order derivative into a single number for encryption, ciphertext transmission, and homomorphic addition operations. The sum of multiple first-order derivatives and second-order derivatives can be simultaneously decoded from the sum of encoded values. We are motivated by the batch idea in the work of BatchCrypt for horizontal federated learning, and design a novel batch method to address the limitations of allowing quite few number of negative numbers. The encode procedure of the proposed batch method consists of four steps, including shifting, truncating, quantizing and batching, while the decoding procedure consists of de-quantization and shifting back. The advantages of our method are demonstrated through theoretical analysis and extensive numerical experiments.

연구 동기 및 목표

  • 수직 분산 XGBoost 학습에서 동형 암호화의 높은 계산 및 통신 오버헤드를 해결하기 위해.
  • 보안 기반의 기울기 집계 과정에서 필요한 암호화 연산 및 암호문 전송 횟수를 줄이기 위해.
  • 수직 분할된 데이터를 가진 기관 간에 효율적이고 개인정보 보호 기반의 XGBoost 학습을 가능하게 하기 위해.
  • 기존의 배치 방법이 기울기 인코딩 중 음수를 처리하는 데 어려움을 겪는 점을 해결하기 위해.
  • SecureBoost와 같은 최신 기법과 비교해 정확도는 유사하게 유지하면서도 효율성을 크게 향상시키기 위해.

제안 방법

  • 일阶 도함수 $g_i$와 이阶 도함수 $h_i$를 4단계 과정(이동, 절삭, 양자화, 배치)을 통해 하나의 수치로 통합 인코딩하기.
  • 인코딩된 단일 값에 대해 동형 암호화 및 덧셈 연산을 수행하여, 동시에 집계된 $g$와 $h$의 계산을 가능하게 하기.
  • 양자화 해제 및 다시 이동하여 인코딩된 값의 합을 복원함으로써 개별 $g$와 $h$ 값을 복원하기.
  • 정밀도 제어 및 인코딩 중 오버플로우 방지를 위해 양자화 파라미터 $\alpha_{\sf{max}}$와 스케일링 인자 $r$을 도입하기.
  • 기존 방법들(예: BatchCrypt)에서 흔히 발생하는 오버플로우 문제를 피하기 위해 음수를 안전하게 처리할 수 있도록 배치 기반 설계를 수행하기.
  • 모델 학습 중 인코딩된 기울기 값에 대해 안전한 덧셈 연산을 수행하기 위해 Paillier 동형 암호화를 사용하기.

실험 결과

연구 질문

  • RQ1일阶 및 이阶 기울기를 배치하여 수직 분산 XGBoost에서 동형 암호화 연산 횟수와 암호문 전송 횟수를 줄일 수 있는가?
  • RQ2기존 방법들과 달리 음수를 처리할 때 오버플로우 없이 어떻게 배치를 설계할 수 있는가?
  • RQ3양자화 파라미터가 기울기 인코딩 시 정밀도 손실과 수치적 안정성에 미치는 영향은 어떠한가?
  • RQ4SecureBoost와 비교해 본다면 제안된 방법은 학습 효율성과 모델 정확도 측면에서 어떤가?
  • RQ5기울기 인코딩에 의한 양자화가 발생하더라도 모델 성능(AUC 및 KS 기준)을 유지할 수 있는가?

주요 결과

  • 제안된 방법은 대규모 데이터셋에서 수직 분산 XGBoost 학습의 총 런타임을 30% 이상 감소시키며, SecureBoost 대비 트리 구축 시간은 최대 50% 감소시킵니다.
  • 91,500개의 샘플을 가진 데이터셋에서 총 런타임은 SecureBoost 대비 75% 수준으로 감소하며, 문제 크기가 커질수록 성능 향상이 더 두드러집니다.
  • 모델 정확도는 유사하게 유지되며, 모든 평가된 데이터셋과 구성에서 AUC 및 KS 값의 차이는 0.005 이내로 매우 미미합니다.
  • 부스팅 라운드 수가 증가할수록 효율성 향상 비율이 증가하여, 90개의 트리에서 트리당 학습 속도가 최대 50% 빨라집니다.
  • $\alpha_{\sf{max}}$와 $r$의 적절한 선택을 통해 근사 최적의 성능를 달성하며, 오버플로우를 최소화하고 정밀도를 유지합니다.
  • 이론적 분석을 통해 음수의 덧셈으로 인한 오버플로우 문제를 피할 수 있음을 확인하였으며, 이는 이전의 배치 기법에서의 핵심 제약 사항을 극복한 것입니다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.