Skip to main content
QUICK REVIEW

[论文解读] Creation and Analysis of an International Corpus of Privacy Laws

Sonu Gupta, Ellen Poplavska|arXiv (Cornell University)|Jun 28, 2022
Privacy, Security, and Data Protection被引用 4
一句话总结

本文介绍了政府隐私指令(GPI)语料库,这是一个大规模、多语言的语料库,包含来自182个司法管辖区的1,043项隐私法律、法规和指南。通过自然语言处理技术,研究揭示了过去50年隐私立法的显著增长,个人数据类型关注度不均,以及金融、医疗保健和电信等反复出现的主题,同时突显了全面隐私法律的稀有性。

ABSTRACT

The landscape of privacy laws and regulations around the world is complex and ever-changing. National and super-national laws, agreements, decrees, and other government-issued rules form a patchwork that companies must follow to operate internationally. To examine the status and evolution of this patchwork, we introduce the Government Privacy Instructions Corpus, or GPI Corpus, of 1,043 privacy laws, regulations, and guidelines, covering 182 jurisdictions. This corpus enables a large-scale quantitative and qualitative examination of legal foci on privacy. We examine the temporal distribution of when GPIs were created and illustrate the dramatic increase in privacy legislation over the past 50 years, although a finer-grained examination reveals that the rate of increase varies depending on the personal data types that GPIs address. Our exploration also demonstrates that most privacy laws respectively address relatively few personal data types, showing that comprehensive privacy legislation remains rare. Additionally, topic modeling results show the prevalence of common themes in GPIs, such as finance, healthcare, and telecommunications. Finally, we release the corpus to the research community to promote further study.

研究动机与目标

  • 为解决国际隐私法律缺乏全面、多语言语料库以支持大规模自然语言处理分析的问题。
  • 研究隐私立法的时间趋势,识别个人数据类型关注焦点的变化。
  • 使用主题建模分析政府隐私指令中的主题模式。
  • 为未来关于法律语言、合规性以及跨司法管辖区隐私监管的研究提供支持。

提出的方法

  • 从182个司法管辖区收集了1,043份政府隐私指令(GPI)语料,包括法律、法规和非约束性指南。
  • 收集原始语言文本及英文翻译,并附带丰富的元数据(URL、颁布日期、司法管辖区、国际协议等)。
  • 应用时间序列分析,追踪1970年至2023年期间GPI的数量与演变。
  • 使用基于LDA的主题建模方法,识别语料库中反复出现的主题。
  • 采用文本分析方法,研究18种个人数据类型在文本中的提及分布。
  • 公开发布该语料库,以支持全球隐私监管的大规模、可复现研究。

实验结果

研究问题

  • RQ1过去50年中,全球隐私立法的数量如何演变?其时间分布中呈现出何种模式?
  • RQ2哪些个人数据类型在隐私法律中被最频繁地提及?这种关注程度在不同司法管辖区和时间点上如何变化?
  • RQ3政府隐私指令中的主导主题聚类是什么?它们与特定行业或数据类型有何关联?
  • RQ4隐私法律在多大程度上共同涵盖多种数据类型?哪些组合最为常见?
  • RQ5隐私法律的语言和结构特征在不同司法管辖区和法律传统之间如何变化?

主要发现

  • 自20世纪70年代以来,隐私法律和法规的数量显著增加,尤其自2000年代起出现急剧上升。
  • 不同个人数据类型的关注度差异显著,某些类型(如金融、健康)受到更多关注。
  • 大多数隐私法律仅涉及少数几种个人数据类型,表明全面隐私立法仍属罕见。
  • 主题建模揭示了反复出现的主题,如金融、医疗保健、电信和生物识别技术,其中生物识别和基因数据常共同出现。
  • 某些数据类型组合表现出强烈的共现模式,包括生物识别与基因数据(直观),以及种族/民族与宗教信仰(较出人意料)。
  • 该语料库揭示了全球隐私监管的碎片化格局,各司法管辖区在范围、语言和关注重点方面存在显著差异。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。