Skip to main content
QUICK REVIEW

[论文解读] Value Engineering for Autonomous Agents

Nieves Montes, Nardine Osman|arXiv (Cornell University)|Feb 17, 2023
Psychology of Moral and Emotional Judgment被引用 7
一句话总结

本文提出了一种新型的人工道德智能体(AMA)范式,将人类选定的价值作为上下文相关的目标嵌入其中,使智能体能够对行为进行规范性推理,并通过具备价值意识的规范评估,集体地使规范与这些价值保持一致。其核心贡献是一个框架,智能体通过将价值与目标结构关联,并利用共识机制共同演化出反映共享人类价值的规范,从而实现价值意识。

ABSTRACT

Machine Ethics (ME) is concerned with the design of Artificial Moral Agents (AMAs), i.e. autonomous agents capable of reasoning and behaving according to moral values. Previous approaches have treated values as labels associated with some actions or states of the world, rather than as integral components of agent reasoning. It is also common to disregard that a value-guided agent operates alongside other value-guided agents in an environment governed by norms, thus omitting the social dimension of AMAs. In this blue sky paper, we propose a new AMA paradigm grounded in moral and social psychology, where values are instilled into agents as context-dependent goals. These goals intricately connect values at individual levels to norms at a collective level by evaluating the outcomes most incentivized by the norms in place. We argue that this type of normative reasoning, where agents are endowed with an understanding of norms' moral implications, leads to value-awareness in autonomous agents. Additionally, this capability paves the way for agents to align the norms enforced in their societies with respect to the human values instilled in them, by complementing the value-based reasoning on norms with agreement mechanisms to help agents collectively agree on the best set of norms that suit their human values. Overall, our agent model goes beyond the treatment of values as inert labels by connecting them to normative reasoning and to the social functionalities needed to integrate value-aware agents into our modern hybrid human-computer societies.

研究动机与目标

  • 解决当前AMA将价值视为静态标签而非推理主动组成部分的缺陷。
  • 通过使智能体能够集体推理并使规范与人类赋予的价值保持一致,整合代理的社会维度。
  • 赋予智能体基于道德与社会心理学而非仅道德哲学的规范性推理能力。
  • 使智能体能够基于其与价值根基目标的对齐程度来评估规范,确保在多智能体系统中的伦理行为。
  • 通过计算机制(如论证与协商)为自主系统中的价值意识协调奠定基础。

提出的方法

  • 价值被形式化为持久的、上下文相关的目标,智能体利用这些目标来指导推理与行动选择。
  • 规范被表示为具有明确语法的规范性规则,由智能体评估其与价值根基目标的对齐程度。
  • 智能体通过分析或模拟来评估规范性结果,以确定其满足价值根基目标的程度。
  • 采用共识机制(如计算社会选择、论证与协商)实现反映系统共享价值的集体规范选择。
  • 智能体架构支持动态的规范创建、选择及基于对结果的道德评估的潜在违规。
  • 该系统不仅支持关于行动的价值意识决策,还支持关于何时应遵守或违反规范的决策。

实验结果

研究问题

  • RQ1如何将价值正式嵌入自主智能体中,使其不仅作为标签,而是作为推理的主动组成部分?
  • RQ2在多智能体环境中,智能体如何集体地使规范与人类赋予的价值保持一致?
  • RQ3何种机制使智能体能够评估规范与价值根基目标之间的道德对齐程度?
  • RQ4智能体如何以反映个体价值与集体规范的方式对规范性行为进行推理?
  • RQ5共识技术在实现多智能体社会中价值对齐的规范性系统中扮演何种角色?

主要发现

  • 所提出的范式使智能体能够通过将价值嵌入持久的、上下文相关的根基目标来实现价值意识,这些目标可指导推理与行动。
  • 智能体能够基于其与价值根基目标的对齐程度来评估规范,采用形式化分析或基于仿真的结果预测。
  • 通过论证与协商等共识机制,促进了集体规范选择,使系统能够收敛到反映共享价值的规范。
  • 该框架支持动态规范性推理,包括基于对结果的道德评估而决定遵守或违反规范。
  • 该方法实现了从将价值视为语法标签到使其成为语义化、基于推理的建构的转变,嵌入于智能体架构之中。
  • 该模型为未来工作中的价值丰富型心智理论奠定了基础,使智能体能够以道德上一致的方式推理他人价值观与行为。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。