Skip to main content
QUICK REVIEW

[논문 리뷰] A Differentially Private Text Perturbation Method Using a Regularized Mahalanobis Metric

Zekun Xu, Abhinav Aggarwal|arXiv (Cornell University)|2020. 10. 22.
Privacy-Preserving Technologies in Data참고 문헌 42인용 수 4
한 줄 요약

이 논문은 단어 임베딩 공분산 구조를 기반으로 소음의 크기를 적응적으로 조정하는 정규화된 마할라노비스 거리 측도를 사용하는 차별적 비공식 텍스트 편향 방법을 제안한다. 이는 유용성 손실 없이 희소 영역에서의 프라이버시를 향상시킨다. 이 방법은 다변량 라플라스 기반 메커니즘과 비교해 더 낮은 단어 비교체 비율과 더 높은 대체 다양성을 보이며, 동일한 하류 분류 성능을 유지한다.

ABSTRACT

Balancing the privacy-utility tradeoff is a crucial requirement of many practical machine learning systems that deal with sensitive customer data. A popular approach for privacy-preserving text analysis is noise injection, in which text data is first mapped into a continuous embedding space, perturbed by sampling a spherical noise from an appropriate distribution, and then projected back to the discrete vocabulary space. While this allows the perturbation to admit the required metric differential privacy, often the utility of downstream tasks modeled on this perturbed data is low because the spherical noise does not account for the variability in the density around different words in the embedding space. In particular, words in a sparse region are likely unchanged even when the noise scale is large. %Using the global sensitivity of the mechanism can potentially add too much noise to the words in the dense regions of the embedding space, causing a high utility loss, whereas using local sensitivity can leak information through the scale of the noise added. In this paper, we propose a text perturbation mechanism based on a carefully designed regularized variant of the Mahalanobis metric to overcome this problem. For any given noise scale, this metric adds an elliptical noise to account for the covariance structure in the embedding space. This heterogeneity in the noise scale along different directions helps ensure that the words in the sparse region have sufficient likelihood of replacement without sacrificing the overall utility. We provide a text-perturbation algorithm based on this metric and formally prove its privacy guarantees. Additionally, we empirically show that our mechanism improves the privacy statistics to achieve the same level of utility as compared to the state-of-the-art Laplace mechanism.

연구 동기 및 목표

  • 텍스트 데이터 편향에서의 프라이버시-유용성 트레이드오프를 해결하기 위해, 특히 임베딩 공간의 저밀도 영역에 있는 단어들에 초점을 맞춘다.
  • 다변량 라플라스 기반 메커니즘에서 구형 소음의 한계를 극복하기 위해, 고밀도 영역에서의 단어를 효과적으로 편향시키지 못하는 문제를 해결한다.
  • 단어 임베딩의 내재된 공분산 구조를 고려한 소음 기반 메커니즘을 설계하여, 희소 영역에서의 대체 가능성 확률을 향상시킨다.
  • 제안된 마할라노비스 기반 편향 기반 메커니즘에 대해 메트릭 차별적 비공식성의 공식적 증명을 수행한다.
  • 실증적으로 이 메커니즘이 유용성과 비교 가능한 최첨단 라플라스 기반 메커니즘과 동일한 프라이버시 예산 하에서 프라이버시 통계를 향상시킴을 검증한다.

제안 방법

  • 텍스트를 연속적인 단어 임베딩으로 매핑하고, 국소 임베딩 공분산에 따라 소음 크기를 조정하기 위해 정규화된 마할라노비스 노름을 사용하는 타원형 소음을 적용한다.
  • 마할라노비스 거리 측도는 어휘의 전체 공분산 행렬에서 유도되며, 고분산 방향으로 소음을 늘여 희소 영역에서의 편향 효과를 높인다.
  • 희소 데이터에 대한 과적합을 방지하고 공분산 추정치의 안정성을 높이기 위해 마할라노비스 거리 측도에 정규화 항을 도입한다.
  • 정규화된 마할라노비스 거리 측도로 정의된 공분산 구조를 가진 다변량 라플라스 분포에서 소음을 샘플링하여, 메트릭 차별적 비공식성을 보장한다.
  • 편향된 임베딩을 근접한 이웃 검색을 통해 이산 어휘 공간으로 다시 투영한다.
  • 프라이버시 예산 ε는 전역 감도를 사용하여 조정되며, 조정 파rameter λ는 정규화 강도를 제어한다.

실험 결과

연구 질문

  • RQ1마할라노비스 기반 소음 기반 메커니즘이 임베딩 공간의 공분산에 적응함으로써 텍스트 편향에서 프라이버시 통계를 향상시킬 수 있는가?
  • RQ2제안된 메커니즘이 다변량 라플라스 기반 메커니즘의 구형 소음과 비교해 희소 임베딩 영역에 있는 단어들의 대체 비율을 더 높일 수 있는가?
  • RQ3마할라노비스 기반 메커니즘이 동일한 프라이버시 예산 하에서 라플라스 기반 메커니즘과 비교해 하류 유용성을 얼마나 잘 유지하는가?
  • RQ4정규화 파rameter λ는 제안된 프레임워크에서 프라이버시와 유용성의 트레이드오프에 어떤 영향을 미치는가?
  • RQ5마할라노비스 기반 메커니즘은 각 단어 클러스터의 국소 공분산 추정으로 확장되어 개인화를 향상시킬 수 있는가?

주요 결과

  • 모든 프라이버시 예산 범위에서 마할라노비스 기반 메커니즘이 평균 비교체 단어 수(Nw)를 크게 감소시키고 평균 고유 대체 수(Sw)를 증가시켰으며, 특히 중간 ε 범위(ε = 5, 10, 20)에서 두드러진다.
  • ε = 10일 때 마할라노비스 기반 메커니즘은 평균 Nw 12.3과 평균 Sw 87.7을 기록했고, 라플라스 기반 메커니즘은 28.1과 71.9를 기록했으며, 95% 신뢰구간을 통한 통계적 유의성 확인이 이루어졌다.
  • 트위터 및 SMSSpam 데이터셋에서, 모든 ε 및 λ 값에서 두 메커니즘 간의 텍스트 분류 정확도, 정밀도, 재현율이 거의 동일하게 유지되어 유용성이 유지됨을 보여주었다.
  • 동일한 소음 스케일에서 다변량 라플라스 기반 메커니즘보다 더 나은 프라이버시 통계를 달성하여, 유용성 손실 없이 더 강력한 프라이버시 보장을 제공함을 시사한다.
  • 300차원의 GloVe 및 FastText 임베딩 양쪽 모두에서 프라이버시 메트릭에서 일관된 향상이 관찰되어, 다양한 임베딩 유형에 대해 메커니즘의 강건성을 입증했다.
  • 정규화 파arameter λ는 프라이버시-유용성 트레이드오프 조정에 효과적이었으며, [0,1] 범위 내에서 그리드 서치로도 충분한 최적화가 가능했다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.