Skip to main content
QUICK REVIEW

[论文解读] Modelling Human Values for AI Reasoning

Nardine Osman, Mark d’Inverno|ArXiv.org|Feb 9, 2024
Explainable Artificial Intelligence (XAI)被引用 4
一句话总结

本文提出了一种形式化、计算基础的人类价值观模型——价值分类模型(VTM),以使人工智能系统能够显式地推理价值观。通过整合社会心理学与形式逻辑的洞见,VTM 捕获了价值观的语义、重要性及其相互关系,使该模型在医疗保健等价值观对齐决策至关重要的领域具备实际应用潜力。

ABSTRACT

One of today's most significant societal challenges is building AI systems whose behaviour, or the behaviour it enables within communities of interacting agents (human and artificial), aligns with human values. To address this challenge, we detail a formal model of human values for their explicit computational representation. To our knowledge, this has not been attempted as yet, which is surprising given the growing volume of research integrating values within AI. Taking as our starting point the wealth of research investigating the nature of human values from social psychology over the last few decades, we set out to provide such a formal model. We show how this model can provide the foundational apparatus for AI-based reasoning over values, and demonstrate its applicability in real-world use cases. We illustrate how our model captures the key ideas from social psychology research and propose a roadmap for future integrated, and interdisciplinary, research into human values in AI. The ability to automatically reason over values not only helps address the value alignment problem but also facilitates the design of AI systems that can support individuals and communities in making more informed, value-aligned decisions. More and more, individuals and organisations are motivated to understand their values more explicitly and explore whether their behaviours and attitudes properly reflect them. Our work on modelling human values will enable AI systems to be designed and deployed to meet this growing need.

研究动机与目标

  • 为解决人工智能中价值观对齐的关键挑战,创建一种形式化、可计算的人类价值观表示方法。
  • 弥合社会心理学中关于价值观的研究与实际人工智能系统设计之间的差距。
  • 使人工智能系统能够基于现实情境,推理人类与人工行为的价值对齐性。
  • 通过人工智能反馈机制,支持个人与组织做出更知情、更具价值观意识的决策。
  • 为未来基于实证社会心理学的跨学科人工智能价值观研究提供基础框架。

提出的方法

  • 利用包含价值观语义、重要性与关系的结构化价值观分类体系,构建人类价值观的形式化模型。
  • 采用形式语言精确表示价值观及其相互依赖关系,确保计算可行性与逻辑一致性。
  • 基于数十年的社会心理学研究构建模型,确保与既定理论在概念上的一致性。
  • 设计模型时保持实现无关性,同时支持具体的数值结构与算法,以实现基于价值观的推理。
  • 通过实际应用验证模型,包括与医疗专业人员合作,评估临床决策中的价值观对齐性。
  • 将形式化模型回溯映射至社会心理学文献,以证明其与价值观结构与动态的实证发现一致。

实验结果

研究问题

  • RQ1如何以一种支持人工智能系统计算推理的方式,形式化表示人类价值观?
  • RQ2为准确反映社会心理学研究,建模价值观语义、重要性与相互关系所需的必要核心组件是什么?
  • RQ3如何将形式化价值观模型实际应用于改善医疗等现实领域中的价值观对齐?
  • RQ4人工智能系统如何利用此类模型,对人类或集体行为的价值对齐性提供反馈?
  • RQ5该模型如何支持个人与组织以结构化方式识别、反思并发展其价值观体系?

主要发现

  • 所提出的价值分类模型(VTM)是首个形式化、可计算的人类价值观模型,明确表示了价值观语义、重要性与关系。
  • 该模型成功捕捉了社会心理学的关键洞见,包括价值观的层级性与关系性,通过与现有研究的详细映射得到验证。
  • VTM 支持实际的人工智能应用,例如在马尔医院评估医疗决策与规程的价值对齐性,从而提升决策支持能力。
  • 该模型支持对随时间推移与跨情境的价值对齐进行推理,可动态评估行为与不断演化的价值观体系之间的关系。
  • 形式化模型为构建不仅与人类价值观对齐,还能帮助用户提升自身价值观意识的人工智能系统提供了基础。
  • 该模型通过实际合作得到验证,证明其在医疗与应急服务等复杂、高价值环境中的可行性与实际影响。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。