Skip to main content
QUICK REVIEW

[論文レビュー] Privacy-Preserving Distributed Learning in the Analog Domain

Mahdi Soleymani, Hessam Mahdavifar|arXiv (Cornell University)|Jul 17, 2020
Cryptography and Data Security被引用数 11
ひとこと要約

本稿は、量子化に起因する精度損失を回避するため、シャミアの秘密分散のアナログ版を用いて、アナログドメインにおける実数および複素数値データのためのプライバシー保護型分散学習フレームワークを提案する。連続的領域における区別安全性(DS)と相互情報量安全性(MIS)の理論的関係を確立し、浮動小数点実装を用いたMNISTにおける固定小数点有限体手法と比較して、優れた精度と効率性を示す。

ABSTRACT

We consider the critical problem of distributed learning over data while keeping it private from the computational servers. The state-of-the-art approaches to this problem rely on quantizing the data into a finite field, so that the cryptographic approaches for secure multiparty computing can then be employed. These approaches, however, can result in substantial accuracy losses due to fixed-point representation of the data and computation overflows. To address these critical issues, we propose a novel algorithm to solve the problem when data is in the analog domain, e.g., the field of real/complex numbers. We characterize the privacy of the data from both information-theoretic and cryptographic perspectives, while establishing a connection between the two notions in the analog domain. More specifically, the well-known connection between the distinguishing security (DS) and the mutual information security (MIS) metrics is extended from the discrete domain to the continues domain. This is then utilized to bound the amount of information about the data leaked to the servers in our protocol, in terms of the DS metric, using well-known results on the capacity of single-input multiple-output (SIMO) channel with correlated noise. It is shown how the proposed framework can be adopted to do computation tasks when data is represented using floating-point numbers. We then show that this leads to a fundamental trade-off between the privacy level of data and accuracy of the result. As an application, we also show how to train a machine learning model while keeping the data as well as the trained model private. Then numerical results are shown for experiments on the MNIST dataset. Furthermore, experimental advantages are shown comparing to fixed-point implementations over finite fields.

研究の動機と目的

  • データが自然に実数/複素数(アナログ)形式である分散学習システムにおけるデータプライバシーの保護という、極めて重要な課題に対処すること。
  • 暗号的セキュリティを確保するための有限体への実数値データの量子化が引き起こす精度の低下を克服すること。
  • 秘密分散の原則を用いて、アナログデータ上で分散計算を安全に実行する情報理論的枠組みを構築すること。
  • 情報理論的および暗号的メトリクスを用いて、アナログドメインにおけるプライバシー保証を定量化すること。
  • 現実世界の学習タスクにおいて、固定小数点代替手法と比較して浮動小数点実装の実現可能性と利点を示すこと。

提案手法

  • 有限体上でのシャミアの秘密分散の連続的対応として、アナログドメインにおける秘密分散方式を導入する。
  • 単一入力複数出力(SIMO)チャネルにおける相関ノイズを用いて、情報漏洩を制限し、区別安全性(DS)の観点からプライバシーを確立する。
  • 離散的領域における区別安全性(DS)と相互情報量安全性(MIS)の既知の関係を、連続的領域へと拡張する。
  • 数値の忠実性を維持し、固定小数点方式で一般的に発生するオーバーフロー問題を回避するため、浮動小数点演算を採用する。
  • データがアナログシェアによって符号化され、並列処理され、個々のシェアが露呈されない形で最終結果が再構成される分散学習プロトコルを設計する。
  • 機械学習モデルの学習—特にMNISTにおけるロジスティック回帰—を、データおよびモデルのプライバシーを損なわずに行えるように、フレームワークを適応させる。

実験結果

リサーチクエスチョン

  • RQ1実数/複素数(アナログ)数の上に、分散学習におけるデータプライバシーを保護する安全な秘密分散方式を構築できるか?
  • RQ2連続的領域における区別安全性(DS)および相互情報量安全性(MIS)といったプライバシーメトリクスを、意味的に拡張し、それらの関係をどのように定式化できるか?
  • RQ3アナログドメインにおける分散学習において、プライバシー水準と計算精度の間の根本的トレードオフは何か?
  • RQ4精度および効率性の観点から、浮動小数点実装は固定小数点有限体方式と比較してどのように性能を発揮するか?
  • RQ5提案されたプロトコルは、データおよびモデルのプライバシーを損なわず、機械学習モデルの学習に効果的に応用可能か?

主な発見

  • 提案されたアナログ秘密分散方式は、連続的領域において情報理論的プライバシー保証を達成し、SIMOチャネル容量の結果を用いて漏洩を上限付ける。
  • 本稿は、アナログドメインにおける区別安全性(DS)と相互情報量安全性(MIS)の理論的同等性を確立し、既知の離散的領域での結果を拡張する。
  • プロトコルの浮動小数点実装は高い精度を維持し、MNISTにおけるロジスティック回帰予測が中央集権ベースラインと密接に一致する。
  • 固定小数点有限体実装と比較して、精度および実行時間の両面で優れた性能を示す。特に、データセットサイズが増加するにつれてその優位性が顕著になる。これは、オーバーフローおよびラップアラウンドエラーを回避できるからである。
  • 符号化計算の通信効率を保ち、MPCベースのアプローチで見られる通信オーバーヘッドを回避する。
  • 本手法はデータセットサイズの増大に対しても頑健であり、大規模な学習においても精度の著しい低下が見られず、固定小数点代替手法がオーバーフローに起因する劣化を受けるのとは対照的である。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。