Skip to main content
QUICK REVIEW

[论文解读] Improving Fairness for Data Valuation in Federated Learning.

Zhenan Fan, Fang Huang|arXiv (Cornell University)|Sep 19, 2021
Privacy-Preserving Technologies in Data参考文献 35被引用 4
一句话总结

本文提出了完整联邦Shapley值,一种增强公平性的联邦学习数据估值方法,通过补全数据拥有者的低秩贡献矩阵,解决了原始联邦Shapley值中的一致性问题。理论与实证结果表明,在温和条件下实现了更高的公平性,确保具有相同数据的拥有者获得相近的估值。

ABSTRACT

Federated learning is an emerging decentralized machine learning scheme that allows multiple data owners to work collaboratively while ensuring data privacy. The success of federated learning depends largely on the participation of data owners. To sustain and encourage data owners' participation, it is crucial to fairly evaluate the quality of the data provided by the data owners and reward them correspondingly. Federated Shapley value, recently proposed by Wang et al. [Federated Learning, 2020], is a measure for data value under the framework of federated learning that satisfies many desired properties for data valuation. However, there are still factors of potential unfairness in the design of federated Shapley value because two data owners with the same local data may not receive the same evaluation. We propose a new measure called completed federated Shapley value to improve the fairness of federated Shapley value. The design depends on completing a matrix consisting of all the possible contributions by different subsets of the data owners. It is shown under mild conditions that this matrix is approximately low-rank by leveraging concepts and tools from optimization. Both theoretical analysis and empirical evaluation verify that the proposed measure does improve fairness in many circumstances.

研究动机与目标

  • 解决原始联邦Shapley值中相同本地数据可能导致不同估值的问题,以提升公平性。
  • 通过使估值独立于数据拥有者排序或子集组成,确保联邦学习中数据估值的公平性。
  • 开发一种保持Shapley值优良性质的同时,提升等价数据贡献间一致性的数据估值方法。
  • 利用优化工具从理论上证明联邦设置下贡献矩阵的低秩结构。
  • 通过实证评估验证所提方法在多样化联邦学习场景中提升数据估值公平性的有效性。

提出的方法

  • 提出一种矩阵补全方法,用于估计联邦学习中数据拥有者所有可能子集的贡献。
  • 采用低秩近似技术建模贡献矩阵,并利用优化理论证明其近似低秩结构的合理性。
  • 应用矩阵补全技术推断贡献矩阵中的缺失条目,确保等价数据集获得一致估值。
  • 通过最小化估值不一致性,推导出完整联邦Shapley值,作为原始联邦Shapley值的更公平替代方案。
  • 通过理论分析表明,在温和条件下,贡献矩阵近似为低秩,从而实现高效且公平的估值估计。
  • 在真实与合成的联邦学习数据集上通过实证评估验证其公平性提升效果。

实验结果

研究问题

  • RQ1矩阵补全技术能否通过减少等价数据贡献的估值不一致性,提升联邦Shapley值的公平性?
  • RQ2在联邦学习设置中,数据拥有者贡献矩阵在何种条件下近似为低秩?
  • RQ3所提出的完整联邦Shapley值是否能确保具有相同本地数据的两个数据拥有者获得相同估值?
  • RQ4在不同联邦学习配置下,完整联邦Shapley值的公平性相较于原始联邦Shapley值有何提升?
  • RQ5所提估值方法的公平性与一致性可提供哪些理论保障?

主要发现

  • 在温和条件下,联邦学习中数据拥有者贡献矩阵近似为低秩,从而支持有效的矩阵补全。
  • 所提完整联邦Shapley值降低了估值不一致性,确保等价数据拥有者获得相近估值。
  • 理论分析证实,矩阵补全方法保留了Shapley值的关键公平性与效率特性。
  • 实证评估表明,该方法在多个联邦学习基准中显著提升了数据估值的公平性。
  • 该方法在保持计算可行性的同时,显著提升了公平性,且未牺牲Shapley值的优良公理性质。
  • 结果表明,当参与者间数据贡献对称或相似时,公平性提升最为显著。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。