[论文解读] Secure Private Information Retrieval from Colluding Databases with Eavesdroppers
本论文研究在N个复制数据库中,面对共谋服务器(最多T个)和窃听者(监听E个数据库)的联合威胁下的私有信息检索(PIR),提出了T-EPIR模型。该模型提出了一种方案,实现下载速率 $ R = \frac{1 - \frac{T}{N}}{1 - \left(\frac{T}{N}\right)^K} - \frac{E}{KN} $,并给出了反向界限,表明当 $ K\to\infty $ 时速率差距趋于零,同时推导出确保安全性的最优共享随机性使用量。
The problem of private information retrieval (PIR) is to retrieve one message out of $K$ messages replicated at $N$ databases, without revealing the identity of the desired message to the databases. We consider the problem of PIR with colluding servers and eavesdroppers, named T-EPIR. Specifically, any $T$ out of $N$ databases may collude, i.e. they may communicate their interactions with the user to guess the identity of the requested message. An eavesdropper is curious to know the database and can tap in on the incoming and outgoing transmissions of any $E$ databases. The databases share some common randomness unknown to the eavesdropper and the user, and use the common randomness to generate the answers, such that the eavesdropper can learn no information about the $K$ messages. Define $R^*$ as the optimal ratio of the number of the desired message information bits to the number of total downloaded bits, and $ρ^*$ to be the optimal ratio of the information bits of the shared common randomness to the information bits of the desired file. In our previous work, we found that when $E \geq T$, the optimal ratio that can be achieved equals $1-\frac{E}{N}$. In this work, we focus on the case when $E \leq T$. We derive an outer bound $R^* \leq (1-\frac{T}{N}) \frac{1-\frac{E}{N} \cdot (\frac{T}{N})^{K-1}}{1-(\frac{T}{N})^K}$. We also obtain a lower bound of $ρ^* \geq \frac{\frac{E}{N}(1-(\frac{T}{N})^K)}{(1-\frac{T}{N})(1-\frac{E}{N} \cdot (\frac{T}{N})^{K-1})}$. For the achievability, we propose a scheme which achieves the rate (inner bound) $R=\frac{1-\frac{T}{N}}{1-(\frac{T}{N})^K}-\frac{E}{KN}$. The amount of shared common randomness used in the achievable scheme is $\frac{\frac{E}{N}(1-(\frac{T}{N})^K)}{1-\frac{T}{N}-\frac{E}{KN}(1-(\frac{T}{N})^K)}$ times the file size. The gap between the derived inner and outer bounds vanishes as the number of messages $K$ tends to infinity.
研究动机与目标
- 建模并分析在存在共谋数据库(最多T个)和窃听者(监听E个数据库)的私有信息检索(PIR)问题。
- 在联合T重共谋与E重窃听条件下,建立安全PIR的信息论极限,特别是当 $ E \leq T $ 时的情况。
- 推导最优检索速率 $ R^* $ 的外边界,以及所需共享随机性比率 $ \rho^* $ 的下界。
- 提出一种新颖的可实现方案,其渐近性能与反向界限一致,当 $ K \to \infty $ 时达到最优。
提出的方法
- 该方案在数据库之间使用共享公共随机性,用户和窃听者均不知晓,以生成隐藏数据库内容的响应。
- 通过在有限域上使用MDS码和随机线性预编码来构建查询与响应,确保用户隐私与系统隐私。
- 在每轮中,用户通过基于文件子集上结构化查询设计的干扰消除机制,从所需文件中检索符号。
- 查询被分配到N个数据库,使得任意T个共谋数据库仅能观察到所有文件的均匀随机线性组合,从而保护用户隐私。
- 窃听者仅能观察到来自共享随机性的EJ个独立符号,确保不会泄露关于数据库的任何信息。
- 通过分析下载总量与所需文件大小的比值,同时考虑共谋与窃听约束,推导出可实现的速率。
实验结果
研究问题
- RQ1当数据库可共谋(最多T个)且窃听者可监听E个数据库时,在 $ E \leq T $ 的约束下,PIR的信息论容量是多少?
- RQ2为确保窃听者无法获取任何关于数据库内容的信息,需要多少共享公共随机性?
- RQ3能否构造一种方案,使其速率接近推导出的反向界限,特别是在消息数量K趋于无穷大时?
- RQ4共谋数据库与窃听者的联合存在如何共同影响PIR速率与保密性要求?
- RQ5在此威胁模型下,下载成本与为实现安全所需共享随机性量之间的权衡是什么?
主要发现
- 最优检索速率的外边界为 $ R^* \leq \left(1 - \frac{T}{N}\right) \frac{1 - \frac{E}{N} \left( \frac{T}{N} \right)^{K-1} }{1 - \left( \frac{T}{N} \right)^K } $,且当 $ K \to \infty $ 时趋于紧致。
- 所提出的方案实现了速率 $ R = \frac{1 - \frac{T}{N}}{1 - \left( \frac{T}{N} \right)^K } - \frac{E}{KN} $,与外边界渐近匹配。
- 所需共享随机性比率满足 $ \rho^* \geq \frac{ \frac{E}{N} \left(1 - \left( \frac{T}{N} \right)^K \right) }{ \left(1 - \frac{T}{N} \right) \left(1 - \frac{E}{N} \left( \frac{T}{N} \right)^{K-1} \right) } $,且该方案实现了该下界。
- 用户隐私得以保护,因为共谋数据库仅能观察到所有文件的均匀随机线性组合,这是由于i.i.d.随机预编码矩阵的作用。
- 系统隐私得到保障,因为窃听者仅能看到共享随机性的独立线性组合,而无法获取实际的数据库内容。
- 内界与外界在检索速率上的差距随着 $ K \to \infty $ 而消失,表明该方案具有渐近最优性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。