[论文解读] Optimizing AI for Teamwork.
本文提出不仅以准确率为唯一目标,而是以人类-AI协作中的人类决定是否接受或覆盖AI建议的团队表现为目标来优化AI系统。通过直接最大化团队效用——平衡决策质量、验证成本及个体准确率——实验表明,在高风险数据集上,团队结果实现了持续且可度量的改进,即使准确率略有下降。
In many high-stakes domains such as criminal justice, finance, and healthcare, AI systems may recommend actions to a human expert responsible for final decisions, a context known as AI-advised decision making. When AI practitioners deploy the most accurate system in these domains, they implicitly assume that the system will function alone in the world. We argue that the most accurate AI team-mate is not necessarily the em best teammate; for example, predictable performance is worth a slight sacrifice in AI accuracy. So, we propose training AI systems in a human-centered manner and directly optimizing for team performance. We study this proposal for a specific type of human-AI team, where the human overseer chooses to accept the AI recommendation or solve the task themselves. To optimize the team performance we maximize the team's expected utility, expressed in terms of quality of the final decision, cost of verifying, and individual accuracies. Our experiments with linear and non-linear models on real-world, high-stakes datasets show that the improvements in utility while being small and varying across datasets and parameters (such as cost of mistake), are real and consistent with our definition of team utility. We discuss the shortcoming of current optimization approaches beyond well-studied loss functions such as log-loss, and encourage future work on human-centered optimization problems motivated by human-AI collaborations.
研究动机与目标
- 解决当前AI系统在人类-AI决策情境中过于注重准确率而忽视团队效能的局限性。
- 探究优化团队表现(而非个体模型准确率)如何在医疗和刑事司法等高风险领域带来更好的实际结果。
- 开发一种直接优化人类-AI团队预期效用的框架,整合决策质量、验证成本及个体表现。
- 证明当AI系统被优化为协作时,即使准确率略有下降,也能在团队效用上实现可度量的提升。
提出的方法
- 将团队效用形式化为决策质量、验证成本以及人类和AI个体准确率的函数。
- 设计一种优化目标,以团队效用为目标,而非标准损失函数(如对数损失)。
- 将此优化方法应用于真实世界、高风险数据集上的线性与非线性模型。
- 使用反映团队层面结果的指标评估性能,包括接受率和最终决策质量。
- 通过建模人类在接受AI建议与独立解决问题之间的选择,将人类决策行为整合进优化过程。
- 通过实证评估比较标准准确率优化模型与团队效用优化模型下的团队效用表现。
实验结果
研究问题
- RQ1将AI系统优化为团队效用(而非个体准确率)是否能带来人类-AI决策表现的可度量改进?
- RQ2在高风险领域中,AI准确率与团队效用之间的权衡如何影响最终决策质量?
- RQ3不同成本结构(如错误成本)在多大程度上影响AI准确率与团队表现之间的最优平衡?
- RQ4人类决定接受或覆盖AI建议的行为如何影响整体团队效用?
- RQ5标准优化方法(如对数损失)在应用于人类-AI协作时存在哪些局限性?
主要发现
- 优化团队效用可在多个真实世界、高风险数据集上实现团队整体表现的一致且可度量的提升。
- 尽管团队效用的提升幅度较小且在不同数据集和参数下有所差异,但其统计显著,且与所提出的团队效用定义一致。
- 当AI被训练为更有效地支持人类决策时,即使准确率略有下降,也能在团队效用上带来显著增益。
- 在验证成本和决策质量至关重要的场景中,所提方法优于标准准确率优化模型。
- 基于对数损失等成熟损失函数的当前优化实践,不足以建模人类-AI协作的动力学。
- 结果表明,未来AI开发应优先考虑以人为本的优化目标,而非纯粹的预测准确率。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。