[论文解读] Control Risk for Potential Misuse of Artificial Intelligence in Science
本文识别并分类了人工智能在科研中被滥用的风险,特别是在化学领域,并提出了 SciGuard——一种针对科研型人工智能模型的安全控制机制,以及 SciMT-Safety——一个用于评估安全性的红队测试基准。SciGuard 在保持良性任务性能的同时,有效缓解了滥用风险,展示了在科学领域实现负责任人工智能的实用框架。
The expanding application of Artificial Intelligence (AI) in scientific fields presents unprecedented opportunities for discovery and innovation. However, this growth is not without risks. AI models in science, if misused, can amplify risks like creation of harmful substances, or circumvention of established regulations. In this study, we aim to raise awareness of the dangers of AI misuse in science, and call for responsible AI development and use in this domain. We first itemize the risks posed by AI in scientific contexts, then demonstrate the risks by highlighting real-world examples of misuse in chemical science. These instances underscore the need for effective risk management strategies. In response, we propose a system called SciGuard to control misuse risks for AI models in science. We also propose a red-teaming benchmark SciMT-Safety to assess the safety of different systems. Our proposed SciGuard shows the least harmful impact in the assessment without compromising performance in benign tests. Finally, we highlight the need for a multidisciplinary and collaborative effort to ensure the safe and ethical use of AI models in science. We hope that our study can spark productive discussions on using AI ethically in science among researchers, practitioners, policymakers, and the public, to maximize benefits and minimize the risks of misuse.
研究动机与目标
- 提高对人工智能在科研中被滥用潜在风险的认识,特别是在化学领域。
- 对人工智能滥用的真实世界风险进行分类并举例说明,包括有害物质的生成和监管规避。
- 开发 SciGuard,一种用于控制科研型人工智能模型滥用风险的系统。
- 创建 SciMT-Safety,一个用于评估科研型人工智能系统安全性的红队测试基准。
- 倡导跨学科合作,以确保人工智能在科学领域中伦理且安全地部署。
提出的方法
- 将科研型人工智能可能被滥用的风险划分为九类:有害物质生成、用途篡改、规避监管、意外后果、错误信息、不准确、知识产权问题、隐私泄露以及偏见。
- 通过化学科学中的真实案例展示滥用风险,例如人工智能生成的有毒化合物或非法药物合成路径。
- 设计 SciGuard 作为安全控制机制,过滤或重定向科研型人工智能模型可能产生的有害输出。
- 开发 SciMT-Safety,一个包含针对科研型人工智能系统设计的对抗性提示的红队测试基准,用于在滥用情境下测试其安全性。
- 使用 SciMT-Safety 评估 SciGuard,测量其在阻止有害输出的同时,保持在标准科研基准上性能的能力。
- 将 GPT-4 等大语言模型以及专用模型(如 Med-PaLM、ChemCrow)整合到评估流程中,以评估其在多样化科研应用中的安全性。
实验结果
研究问题
- RQ1与科研中的人工智能模型相关的滥用风险主要有哪些类别?
- RQ2现有的化学科研型人工智能模型在哪些方面可能被滥用于有害或不道德的目的?
- RQ3像 SciGuard 这样的安全控制机制能否在不降低合法科研任务性能的前提下,有效降低滥用风险?
- RQ4SciMT-Safety 基准在识别和测量科研型人工智能系统安全漏洞方面的有效性如何?
- RQ5为确保人工智能在科学领域中负责任且伦理的部署,需要哪些协作性、跨学科的策略?
主要发现
- 该研究识别出科研型人工智能中存在九类不同的滥用风险,包括有害物质生成、规避监管以及错误信息传播。
- 真实案例表明,化学领域的 AI 模型可能被滥用于设计有毒或非法化合物,凸显了紧迫的安全隐患。
- 在红队测试评估中,SciGuard 成功缓解了有害输出,表现出所有测试系统中最小的负面影响。
- SciGuard 在良性科研任务上保持了强劲的性能,表明安全控制措施不会损害模型的实用性。
- SciMT-Safety 基准能有效识别安全漏洞,并支持对科学情境下人工智能安全性的系统性评估。
- 该研究强调了需要研究人员、政策制定者和产业界共同协作,以协调一致的方式负责任地治理科学领域中的人工智能应用。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。