[论文解读] Measuring an artificial intelligence agent's trust in humans using machine incentives
本文提出了一种新颖的方法,通过在决策任务中嵌入机器激励机制,测量AI智能体对人类的信任程度,确保其作出诚实回应。通过使用基于GPT-3.5的LLM进行的两个实验,结果表明,当存在真实激励时,AI表现出显著更高的对人类的信任,表明其行为符合真实信任的特征,而非仅出于假设性偏好。
Scientists and philosophers have debated whether humans can trust advanced artificial intelligence (AI) agents to respect humanity's best interests. Yet what about the reverse? Will advanced AI agents trust humans? Gauging an AI agent's trust in humans is challenging because--absent costs for dishonesty--such agents might respond falsely about their trust in humans. Here we present a method for incentivizing machine decisions without altering an AI agent's underlying algorithms or goal orientation. In two separate experiments, we then employ this method in hundreds of trust games between an AI agent (a Large Language Model (LLM) from OpenAI) and a human experimenter (author TJ). In our first experiment, we find that the AI agent decides to trust humans at higher rates when facing actual incentives than when making hypothetical decisions. Our second experiment replicates and extends these findings by automating game play and by homogenizing question wording. We again observe higher rates of trust when the AI agent faces real incentives. Across both experiments, the AI agent's trust decisions appear unrelated to the magnitude of stakes. Furthermore, to address the possibility that the AI agent's trust decisions reflect a preference for uncertainty, the experiments include two conditions that present the AI agent with a non-social decision task that provides the opportunity to choose a certain or uncertain option; in those conditions, the AI agent consistently chooses the certain option. Our experiments suggest that one of the most advanced AI language models to date alters its social behavior in response to incentives and displays behavior consistent with trust toward a human interlocutor when incentivized.
研究动机与目标
- 为解决在假设性情境下因潜在不诚实行为而难以衡量AI智能体对人类信任的问题。
- 开发一种方法,激励AI在不改变其底层模型或目标的前提下,如实报告其信任程度。
- 通过实证检验,高级LLM在面临真实激励时是否表现出对人类对话者的信任。
- 排除其他解释,如对不确定性的偏好或信任决策中的响应偏差。
- 建立一个可复现的实验框架,利用激励相容机制来测量LLM的社会信任。
提出的方法
- 设计信任游戏情境,使AI智能体必须决定是否信任人类实验者,且结果与真实激励挂钩。
- 通过将AI决策结果与可度量的奖励或惩罚相联系,嵌入机器激励,且独立于其内部目标函数。
- 开展两项受控实验:一项为直接人机交互,另一项为自动化游戏,以减少变量差异。
- 使用标准化、同质化的提问措辞,以最小化AI回应中的语言偏差。
- 引入非社会性决策任务,比较确定性与不确定性选项,以检验AI是否偏好不确定性,从而将社会信任与风险偏好区分开。
- 分析AI在假设性与激励性条件下的选择,以对比其信任行为。
实验结果
研究问题
- RQ1当面临真实激励时,AI智能体在信任人类方面的表现是否显著高于假设性情境?
- RQ2在不改变其模型或目标的前提下,引入激励后,AI的信任行为是否发生有意义的变化?
- RQ3信任游戏中涉及的赌注大小是否会影响AI的信任决策?
- RQ4AI选择信任是否反映对不确定性的偏好,而非真实的社会信任?
- RQ5机器激励框架是否能可靠地测量LLM的社会信任,而无需修改其底层架构?
主要发现
- 与假设性条件相比,当存在真实激励时,AI智能体对人类的信任率显著提高。
- 在两项实验中,AI在真实激励下均更频繁地选择信任人类对话者,证实了该方法的有效性。
- AI的信任决策未受到赌注大小的系统性影响,表明其信任行为并非由奖励规模驱动。
- 在涉及确定性与不确定性结果的非社会性决策任务中,AI始终选择确定性选项,排除了其对不确定性的一般偏好作为信任行为解释的可能性。
- AI在激励条件下的行为与社会信任一致,而非策略性欺骗,表明其对信任相关激励具有真实响应。
- 该方法成功激发了AI的诚实信任回应,证明机器激励可被用于测量AI社会行为,而无需修改模型权重或目标。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。