[论文解读] Reducing malicious use of synthetic media research: Considerations and potential release practices for machine learning
本文提出了一套细致的框架,用于管理机器学习研究在合成媒体领域的发布,以减少恶意使用,倡导基于情境的、由专家指导的发布实践,而非简单的开放或保密。它概述了风险缓解策略,推动社区规范进行影响评估,并建议机构支持负责任的研究传播,以在创新与社会安全之间取得平衡。
The aim of this paper is to facilitate nuanced discussion around research norms and practices to mitigate the harmful impacts of advances in machine learning (ML). We focus particularly on the use of ML to create "synthetic media" (e.g. to generate or manipulate audio, video, images, and text), and the question of what publication and release processes around such research might look like, though many of the considerations discussed will apply to ML research more broadly. We are not arguing for any specific approach on when or how research should be distributed, but instead try to lay out some useful tools, analogies, and options for thinking about these issues. We begin with some background on the idea that ML research might be misused in harmful ways, and why advances in synthetic media, in particular, are raising concerns. We then outline in more detail some of the different paths to harm from ML research, before reviewing research risk mitigation strategies in other fields and identifying components that seem most worth emulating in the ML and synthetic media research communities. Next, we outline some important dimensions of disagreement on these issues which risk polarizing conversations. Finally, we conclude with recommendations, suggesting that the machine learning community might benefit from: working with subject matter experts to increase understanding of the risk landscape and possible mitigation strategies; building a community and norms around understanding the impacts of ML research, e.g. through regular workshops at major conferences; and establishing institutions and systems to support release practices that would otherwise be onerous and error-prone.
研究动机与目标
- 解决人们对通过机器学习生成的合成媒体(尤其是深度伪造和AI生成的虚假信息)被恶意使用的日益增长的担忧。
- 通过基于风险分析的、涵盖多种发布实践的谱系,减少机器学习社区在开放与限制研究发布之间的两极分化。
- 支持发展社区规范和制度化机制,以实现负责任的研究传播,同时不损害安全或创新。
- 鼓励机器学习研究人员与虚假信息、安全和政策等领域的主题专家合作,以更好地理解并缓解现实世界中的危害。
- 建立结构化、可重复的过程,以评估和管理发布潜在有害机器学习研究的风险,特别是在合成媒体领域。
提出的方法
- 基于潜在滥用路径,建立研究危害的分类体系——包括产品危害、数据危害和注意力危害。
- 将其他高风险领域(如生物技术和核科学中的双重用途研究)的风险缓解策略,适配至机器学习情境。
- 提出超越二元开放/封闭模型的多种发布选项,包括延迟发布、受限访问和内容删减出版。
- 倡导对研究提案进行专家影响评估,引入领域专家以评估风险和缓解潜力。
- 设计原型审查系统,以安全方式共享敏感模型,减少对研究人员个人进行临时验证的依赖。
- 建立定期举办的研讨会和社区论坛,以制度化影响评估,并促进负责任研究实践的共享规范。
实验结果
研究问题
- RQ1如何以最小化恶意使用风险的方式发布关于合成媒体的机器学习研究,同时保持科学开放性?
- RQ2与合成媒体研究相关的关键危害类型(产品、数据、注意力)有哪些?它们在风险特征上如何不同?
- RQ3机器学习社区如何发展共享规范和制度化结构,以评估和管理机器学习研究的社会风险?
- RQ4其他科学领域中双重用途研究的经验,对管理合成媒体机器学习研究中的风险有何启示?
- RQ5哪些发布实践可以在透明度、创新和社会安全之间实现平衡?
主要发现
- 本文识别出三种截然不同的危害类型——产品危害、数据危害和注意力危害——有助于对机器学习在合成媒体研究中可能被恶意利用的方式进行分类。
- 研究表明,研究发布的争论并非简单的开放与封闭访问之间的二元对立,而是涉及一系列基于情境的选项。
- 作者发现,专家影响评估和制度化支持可显著减轻个体研究人员的负担,同时提高风险评估的准确性。
- 他们表明,当前对敏感模型请求者的临时验证方法存在易错性且不可持续,亟需可扩展的审查系统。
- 本文结论认为,负责任的发布实践是可行且必要的,制度化影响评估有助于机器学习社区在社会责任方面实现成熟。
- 它强调,与受影响社区及主题专家的主动协作至关重要,以避免风险夸大,确保风险映射的现实性和可操作性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。