Skip to main content
QUICK REVIEW

[论文解读] Requisite Variety in Ethical Utility Functions for AI Value Alignment

Nadisha-Marie Aliman, Leon Kester|arXiv (Cornell University)|Jun 30, 2019
Psychology of Moral and Emotional Judgment参考文献 32被引用 10
一句话总结

本文提出将神经科学与心理学中的科学洞见整合到人工智能的伦理效用函数中,以确保其能够捕捉人类道德直觉的全部多样性,采用增强功利主义作为非规范性伦理框架。该研究引入一种社会技术反馈回路,通过模拟未来体验来验证这些效用函数,旨在增强对对抗性攻击的鲁棒性,并改善人工智能的价值对齐。

ABSTRACT

Being a complex subject of major importance in AI Safety research, value alignment has been studied from various perspectives in the last years. However, no final consensus on the design of ethical utility functions facilitating AI value alignment has been achieved yet. Given the urgency to identify systematic solutions, we postulate that it might be useful to start with the simple fact that for the utility function of an AI not to violate human ethical intuitions, it trivially has to be a model of these intuitions and reflect their variety $ - $ whereby the most accurate models pertaining to human entities being biological organisms equipped with a brain constructing concepts like moral judgements, are scientific models. Thus, in order to better assess the variety of human morality, we perform a transdisciplinary analysis applying a security mindset to the issue and summarizing variety-relevant background knowledge from neuroscience and psychology. We complement this information by linking it to augmented utilitarianism as a suitable ethical framework. Based on that, we propose first practical guidelines for the design of approximate ethical goal functions that might better capture the variety of human moral judgements. Finally, we conclude and address future possible challenges.

研究动机与目标

  • 解决在设计与人类道德直觉对齐的伦理效用函数时缺乏共识的问题。
  • 将人工智能价值对齐重新定义为一个安全问题,要求对效用函数的对抗性操纵具备韧性。
  • 将神经科学与心理学中关于具身道德——尤其是情感与二元认知——的科学知识整合到效用函数设计中。
  • 制定实用指南,用于构建反映人类道德判断所需多样性的伦理目标函数。
  • 提出一种基于模拟未来体验(如VR/AR)的验证框架,以评估社会层面的效用与福祉结果。

提出的方法

  • 应用良好调节者定理,论证伦理效用函数必须建模人类伦理直觉,以避免违反伦理直觉。
  • 利用建构主义神经科学与认知科学,将情感与情绪建模为道德判断的核心,强调其具身性。
  • 引入二元道德作为理解道德判断如何从社会与感知语境中产生的心理学框架。
  • 采用增强功利主义——一种非规范性、基于科学的伦理框架——作为制定伦理目标函数的基础。
  • 将社会层面的伦理目标函数 $U_{Total}(s,a,s^{ ext{′}})$ 定义为基于共识的感知者依赖型效用加权聚合,使用共识参数集。
  • 提出通过预演模拟来验证效用函数,即让具有代表性的社会群体通过沉浸式媒体(如VR、音频故事)体验未来情景,利用核心情感效价的时间积分测量人工模拟的未来即时效用 $U_{TotalAS}$:$U_{TotalAS}(s,a,s^{ ext{′}}) \approx \sum_{n=1}^{N} \int_{t_0}^{T} I_n(t)dt$。

实验结果

研究问题

  • RQ1如何设计伦理效用函数,以捕捉人类道德直觉所需的多样性,而不过度依赖学习方法?
  • RQ2情感与二元感知在多大程度上塑造道德判断,这些因素如何在AI效用函数中建模?
  • RQ3如何利用非规范性伦理框架(如增强功利主义)来构建感知者依赖型伦理目标函数?
  • RQ4模拟社会体验(如VR、AR)在部署前验证伦理效用函数方面发挥什么作用?
  • RQ5社会技术反馈回路如何确保伦理目标函数随时间推移持续验证与动态优化?

主要发现

  • 根据良好调节者定理,伦理效用函数必须建模人类道德直觉的多样性,以避免违反伦理直觉。
  • 根植于具身认知的情感与二元认知成分必须整合进效用函数,以确保对对抗性操纵的鲁棒性。
  • 所提出的社会层面伦理目标函数 $U_{Total}(s,a,s^{ ext{′}})$ 定义为基于共识的感知者依赖型效用聚合,实现集体伦理对齐。
  • 人工模拟的未来即时效用 $U_{TotalAS}$ 可通过沉浸式模拟中核心情感效价 $I_n(t)$ 的时间积分进行测量,作为未来福祉的代理指标。
  • 通过模拟社会体验进行部署前验证,可实现对伦理对齐的早期评估,降低真实世界部署前的风险。
  • 通过预先商定的指标(如社会满意度或二元感知度)进行部署后验证,可实现对伦理目标函数的持续监控与动态调整。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。