[Paper Review] A computational framework for human values
This paper proposes a formal computational framework for human values grounded in social science research, enabling AI systems to represent, reason about, and align with human values through a structured value taxonomy. The framework supports dynamic value modeling, value alignment measurement, and integrates with multi-agent systems, offering a foundational model for ethical AI design with provable alignment to human values.
In the diverse array of work investigating the nature of human values from psychology, philosophy and social sciences, there is a clear consensus that values guide behaviour. More recently, a recognition that values provide a means to engineer ethical AI has emerged. Indeed, Stuart Russell proposed shifting AI's focus away from simply ``intelligence'' towards intelligence ``provably aligned with human values''. This challenge -- the value alignment problem -- with others including an AI's learning of human values, aggregating individual values to groups, and designing computational mechanisms to reason over values, has energised a sustained research effort. Despite this, no formal, computational definition of values has yet been proposed. We address this through a formal conceptual framework rooted in the social sciences, that provides a foundation for the systematic, integrated and interdisciplinary investigation into how human values can support designing ethical AI.
Motivation & Objective
- To address the lack of a formal, computational definition of human values in AI research.
- To provide a unified, interdisciplinary foundation for modeling human values that supports ethical AI development.
- To enable systematic reasoning over individual and group values, including dynamic changes over time.
- To formalize the value alignment problem, allowing for provable alignment between AI behavior and human values.
- To support the design of data structures and algorithms for value-aware AI systems.
Proposed method
- Proposes a formal value taxonomy using abstract concepts (labels), property nodes for meaning, and importance weights to represent value coherence and hierarchy.
- Introduces a computational model of value alignment based on the degree of behavioral consistency with selected values, using a normalized similarity measure.
- Models individual and group values through hybrid multi-agent systems, integrating artificial and human agents in organizational or community settings.
- Incorporates contextual evaluation to model dynamic value evolution, where agents re-evaluate values based on situational contexts.
- Uses a running example to demonstrate implementation choices, illustrating value representation, alignment computation, and context-dependent reasoning.
- Employs a formal language to ensure precision, support algorithm development, and enable proof-based reasoning over values.
Experimental results
Research questions
- RQ1How can human values be formally defined and represented in a way that supports computational reasoning in AI?
- RQ2What mechanisms enable the dynamic modeling of individual and group values over time?
- RQ3How can value alignment between AI behavior and human values be formally measured and quantified?
- RQ4What are the computational and representational requirements for aggregating values across individuals and groups?
- RQ5How can value taxonomies be constructed and learned from human stakeholders, both explicitly and implicitly?
Key findings
- The proposed framework provides a formal, computationally tractable model of human values that integrates core principles from social science research, including value coherence and importance weighting.
- Value alignment is formally defined as a normalized similarity measure between observed behavior and selected values, with a computed value of 0.475 in the running example.
- The model supports dynamic value reasoning through context-sensitive evaluation, enabling agents to adapt their value-based decisions over time.
- The framework enables systematic algorithm development for value alignment and aggregation, offering a foundation for provably aligned AI systems.
- The approach is extensible to multi-agent systems and supports both explicit and implicit value acquisition from human stakeholders.
- The model’s design allows for integration with existing multi-agent systems (MAS) concepts while maintaining theoretical clarity and implementation flexibility.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.