Skip to main content
QUICK REVIEW

[论文解读] Reclaiming the Digital Commons: A Public Data Trust for Training Data

Alan Chan, Herbie Bradley|arXiv (Cornell University)|Mar 16, 2023
Privacy-Preserving Technologies in Data被引用 5
一句话总结

本文提出建立一个国家级公共数据信托,以管理基础模型的训练数据,通过将网络抓取的数据授权给商业开发人员并获取收益分成,实现对人工智能发展的集体控制。该信托旨在重新分配经济价值,减少负面外部性,并通过验证、激励和监管协同实现民主监督。

ABSTRACT

Democratization of AI means not only that people can freely use AI, but also that people can collectively decide how AI is to be used. In particular, collective decision-making power is required to redress the negative externalities from the development of increasingly advanced AI systems, including degradation of the digital commons and unemployment from automation. The rapid pace of AI development and deployment currently leaves little room for this power. Monopolized in the hands of private corporations, the development of the most capable foundation models has proceeded largely without public input. There is currently no implemented mechanism for ensuring that the economic value generated by such models is redistributed to account for their negative externalities. The citizens that have generated the data necessary to train models do not have input on how their data are to be used. In this work, we propose that a public data trust assert control over training data for foundation models. In particular, this trust should scrape the internet as a digital commons, to license to commercial model developers for a percentage cut of revenues from deployment. First, we argue in detail for the existence of such a trust. We also discuss feasibility and potential risks. Second, we detail a number of ways for a data trust to incentivize model developers to use training data only from the trust. We propose a mix of verification mechanisms, potential regulatory action, and positive incentives. We conclude by highlighting other potential benefits of our proposed data trust and connecting our work to ongoing efforts in data and compute governance.

研究动机与目标

  • 解决私人企业对人工智能发展和训练数据的权力过度集中问题。
  • 纠正人工智能带来的负面外部性,如数字公域的退化和经济不平等。
  • 建立机制,将人工智能生成的经济价值重新分配给数据贡献者。
  • 使公民能够就人工智能系统的训练和部署方式共同行使决策权。
  • 创建一个切实可行、可扩展的训练数据治理模型,以支持公共利益和民主监督。

提出的方法

  • 建立国家级公共数据信托,收集并授权从公共互联网抓取的预训练数据以及标注者提供的人工反馈数据。
  • 实施收益分成模式,即商业模型开发者向信托支付其部署收入的一定比例,以换取使用其数据的权限。
  • 通过密码学证明、数据溯源追踪和第三方审计建立验证机制,确保开发者仅使用信托授权的数据。
  • 结合监管激励(如符合数据治理法律)、正向激励(如优先访问权、认证资格)和技术强制手段,推动广泛采用。
  • 将信托整合进更广泛的AI治理框架中,与现有的数据 stewardship(数据治理)和计算治理努力保持一致。
  • 将信托定位为公共产品提供者,支持可持续的数据生成,保障长期数字公域的存续。

实验结果

研究问题

  • RQ1如何通过公共数据信托从私人企业手中夺回对训练数据的控制权,并恢复民主监督?
  • RQ2哪些机制可确保模型开发者仅使用信托授权的数据,以及如何验证其合规性?
  • RQ3公共数据信托如何实现可持续的资金来源,并确保人工智能生成价值的公平再分配?
  • RQ4国家级公共数据信托在缓解人工智能部署负面外部性方面可发挥何种作用?
  • RQ5公共数据信托如何与现有数据治理和计算治理框架共存并形成互补?

主要发现

  • 公共数据信托可作为一项切实可行的制度机制,重新夺回数字公域控制权,并重新分配人工智能发展带来的经济价值。
  • 诸如密码学证明和数据溯源追踪等验证机制,可有效确保模型开发者仅使用信托授权的数据。
  • 收益分成模式为人工智能部署利润向公众再分配提供了直接渠道,有助于纠正外部性并促进正义。
  • 该信托模型通过监管协同、认证优势和优质数据的优先访问权,激励开发者合规。
  • 该信托模式支持将训练数据生成视为公共产品,降低数字公域退化的风险,提升长期可持续性。
  • 该信托与现有治理框架(尤其在数据治理和计算治理方面)相辅相成,为国家人工智能政策提供可扩展的治理范式。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。