[论文解读] Robust Privatization with Multiple Tasks and the Optimal Privacy-Utility Tradeoff
该论文提出了一种针对非特定任务的数据发布鲁棒私有化框架,在满足多个可能任务的效用约束的同时,最小化隐私泄露。通过将问题分解为带权重分量的并行隐私漏斗问题,推导出确定性私有特征的闭式解,并表明最优权重可通过线性规划求得,从而在足够高的发布速率下实现最小隐私泄露。
In this work, fundamental limits and optimal mechanisms of privacy-preserving data release that aims to minimize the privacy leakage under utility constraints of a set of multiple tasks are investigated. While the private feature to be protected is typically determined and known by the sanitizer, the target task is usually unknown. To address the lack of information on the specific task, utility constraints laid on a set of multiple possible tasks are considered. The mechanism protects the specific privacy feature of the to-be-released data while satisfying utility constraints of all possible tasks in the set. First, the single-letter characterization of the rate-leakage-distortion region is derived, where the utility of each task is measured by a distortion function. It turns out that the minimum privacy leakage problem with log-loss distortion constraints and the unconstrained released rate is a non-convex optimization problem. Second, focusing on the case where the raw data consists of multiple independent components, we show that the above non-convex optimization problem can be decomposed into multiple parallel privacy funnel (PF) problems with different weightings. We explicitly derive the optimal solution to each PF problem when the private feature is a component-wise deterministic function of a data vector. The solution is characterized by a leakage-free threshold: when the utility constraint is below the threshold, the minimum leakage is zero; once the required utility level is above the threshold, the privacy leakage increases linearly. Finally, we show that the optimal weighting of each privacy funnel problem can be found by solving a linear program (LP). A sufficient released rate to achieve the minimum leakage is also derived. Numerical results are shown to illustrate the robustness of our approach against the task non-specificity.
研究动机与目标
- 解决在下游使用发布数据的任务未知时保护隐私的挑战。
- 在一组可能任务的效用约束下最小化隐私泄露,确保对任务非特定性的鲁棒性。
- 在信息论框架下,以对数损失失真度量,推导隐私-效用权衡的根本极限。
- 表明最优解可通过将问题分解为带权重分量的并行隐私漏斗问题获得。
- 建立实现最小隐私泄露所需的足够发布速率,并提供最优权重的线性规划公式。
提出的方法
- 论文采用对数损失失真作为效用度量,互信息作为隐私度量,构建隐私-效用权衡的模型。
- 推导出多个效用约束下速率-泄露-失真区域的单字母表征。
- 在数据分量独立且私有特征为确定性的假设下,问题被分解为具有不同加权的并行隐私漏斗问题。
- 每个隐私漏斗问题均以闭式求解,揭示了在泄漏自由阈值以下隐私泄露为零;超过该阈值后,隐私泄露随所需效用水平线性增加。
- 通过求解线性规划(LP)确定每个分量的最优权重,从而实现最小泄露的高效计算。
- 推导出实现最小泄露所需的足够发布速率,确保最优解的可行性。

实验结果
研究问题
- RQ1当下游任务未知且存在多个可能任务时,隐私-效用权衡的根本极限是什么?
- RQ2如何在一组可能任务的效用约束下最小化隐私泄露?
- RQ3能否将最小隐私泄露的非凸优化问题分解为更简单的子问题?
- RQ4在多个效用约束下,数据分量在私有化机制中的最优加权是什么,以最小化泄露?
- RQ5在鲁棒设置下,实现最优隐私-效用权衡所需的足够发布速率是多少?
主要发现
- 在多个效用约束下,最小隐私泄露问题为非凸问题,但在独立性与确定性私有特征假设下,可分解为并行隐私漏斗问题。
- 对于每个分量,当低于泄漏自由阈值时,隐私泄露为零;超过该阈值后,隐私泄露随所需效用水平线性增加。
- 每个隐私漏斗分量的最优权重可通过线性规划计算,实现高效优化。
- 推导出实现最小泄露所需的足够发布速率,并在采用并行私有化时证明其紧致性。
- 数值结果验证了该方法对任务非特定性的鲁棒性,并展示了任务集选择对隐私-效用权衡的影响。
- 该框架可扩展至差分隐私,其中最小泄露问题同样分解为并行的单约束问题,但闭式解仍为开放问题。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。