Skip to main content
QUICK REVIEW

[论文解读] The tradeoff between the utility and risk of location data and implications for public good

Dan Calacci, Alex Berke|arXiv (Cornell University)|May 22, 2019
Human Mobility and Location-Based Analysis参考文献 33被引用 8
一句话总结

本文提出了一种权衡曲线框架,以平衡高分辨率位置数据在效用与隐私风险之间的关系,认为尽管此类数据在城市规划和交通建模等方面具有显著的公共利益潜力,但当前私人市场中的使用方式却损害了伦理与隐私关切。作者主张采用情境感知的数据共享模式,通过负责任的聚合与治理,优先实现公共利益。

ABSTRACT

High-resolution individual geolocation data passively collected from mobile phones is increasingly sold in private markets and shared with researchers. This data poses significant security, privacy, and ethical risks: it's been shown that users can be re-identified in such datasets, and its collection rarely involves their full consent or knowledge. This data is valuable to private firms (e.g. targeted marketing) but also presents clear value as a public good. Recent public interest research has demonstrated that high-resolution location data can more accurately measure segregation in cities and provide inexpensive transit modeling. But as data is aggregated to mitigate its re-identifiability risk, its value as a good diminishes. How do we rectify the clear security and safety risks of this data, its high market value, and its potential as a resource for public good? We extend the recently proposed concept of a tradeoff curve that illustrates the relationship between dataset utility and privacy. We then hypothesize how this tradeoff differs between private market use and its potential use for public good. We further provide real-world examples of how high resolution location data, aggregated to varying degrees of privacy protection, can be used in the public sphere and how it is currently used by private firms.

研究动机与目标

  • 分析个人层面位置数据的高市场价值与其对隐私和安全造成的风险之间的张力。
  • 挑战现有计算度量方法中忽略数据使用中社会与情境因素的效用与风险指标。
  • 探讨聚合水平如何影响公共与私人应用中数据的效用与再识别风险。
  • 提出向以公共利益为优先的数据共享模式进行规范性转变,特别是在通过公共基础设施收集数据的情况下。
  • 通过分析私营企业与公共机构在现实世界中使用位置数据的案例,凸显数据获取与问责机制之间的差异。

提出的方法

  • 将数据效用与隐私风险之间的权衡曲线概念扩展至高分辨率位置数据,进行具体应用。
  • 使用轨迹数据——带时间戳的GPS坐标——来建模个人移动模式,并推断家庭位置与出行方式等行为。
  • 分析来自LBS企业的实际数据集,重点关注不同聚合层级如何在保留分析价值的同时降低再识别风险。
  • 比较私人市场用途(如定向广告)与公共利益应用(如测量城市隔离、交通建模)之间的差异。
  • 借鉴现有公共数据共享模式(如城市自行车共享、网约车数据)的先例,提出针对LBS数据的监管与制度框架。
  • 提出去中心化的数据所有权模型,如Safe Answers,以重构数据治理并确保公共问责性。

实验结果

研究问题

  • RQ1当位置数据被聚合以降低再识别风险时,其效用如何变化?
  • RQ2当前私营市场对位置数据的使用在哪些方面与公共利益应用相冲突?
  • RQ3在未获得知情同意的情况下出售伪匿名化位置数据,其伦理与隐私风险是什么?
  • RQ4如何设计数据共享框架,以确保从公众收集的数据真正惠及公众?
  • RQ5哪些制度或监管模式能够实现位置数据在公共利益中的公平获取,同时将损害最小化?

主要发现

  • 高分辨率位置数据可用于准确测量城市隔离并建模公共交通模式,展现出强大的公共利益潜力。
  • 将位置数据聚合至普查区或统计区域级别可降低再识别风险,但会削弱其在精细化城市研究中的分析效用。
  • 尽管通过应用程序和SDK广泛收集位置数据,用户通常对其收集范围及共享方式缺乏认知,引发严重的知情同意与透明度问题。
  • 私营企业如数据经纪商和LBS提供商通过出售位置数据获利,却未与公共机构共享,导致数据获取与价值捕获的不平衡。
  • 当私营企业使用公共基础设施(如城市道路、补贴的电信塔)时,已有公共数据共享的先例,表明LBS数据也可能适用类似义务。
  • 当前数据市场缺乏监管,若无制度性改革,即使私营企业从公众数据中提取价值,个人风险仍将持续存在。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。