[论文解读] Uncalibrated Models Can Improve Human-AI Collaboration
本文提出,故意使AI置信度得分高于其实际准确率(即人为地使置信度更高),可提升人机协作的表现。通过从数千次人机交互中学习人类行为,作者优化了一种AI置信度的单调变换,以最大化人类的准确率与置信度,结果在四个不同任务中均展现出显著提升,且实验对象为真实用户。
In many practical applications of AI, an AI model is used as a decision aid for human users. The AI provides advice that a human (sometimes) incorporates into their decision-making process. The AI advice is often presented with some measure of "confidence" that the human can use to calibrate how much they depend on or trust the advice. In this paper, we present an initial exploration that suggests showing AI models as more confident than they actually are, even when the original AI is well-calibrated, can improve human-AI performance (measured as the accuracy and confidence of the human's final prediction after seeing the AI advice). We first train a model to predict human incorporation of AI advice using data from thousands of human-AI interactions. This enables us to explicitly estimate how to transform the AI's prediction confidence, making the AI uncalibrated, in order to improve the final human prediction. We empirically validate our results across four different tasks--dealing with images, text and tabular data--involving hundreds of human participants. We further support our findings with simulation analysis. Our findings suggest the importance of jointly optimizing the human-AI system as opposed to the standard paradigm of optimizing the AI model alone.
研究动机与目标
- 探究是否可通过为人类使用而优化AI置信度(而非为模型校准)来改善协作决策表现。
- 解决当前AI系统仅孤立优化模型、而未考虑人类决策偏差的问题。
- 开发一种基于实证的人机交互数据的框架,以转换AI置信度得分,从而更好地支持人类判断。
- 验证未经校准的AI建议是否可带来比校准建议更高的准确率与更优的置信度校准。
提出的方法
- 作者在四个任务(图像、文本、表格)中收集了数千次人机交互数据,以建模用户如何采纳AI建议。
- 训练一个人类行为模型,以预测用户如何根据AI建议调整其置信度与决策。
- 利用该模型,优化AI预测置信度的单调变换,以最大化最终的人类准确率。
- 该变换有意使AI显得比实际更自信,从而生成‘人机校准’的建议。
- 该方法通过模拟验证,随后在Prolific平台上通过众包研究对真实人类参与者进行测试。
- 该方法进一步在另一批英国参与者中进行测试,以评估其跨文化泛化能力。
实验结果
研究问题
- RQ1是否可通过故意未校准AI置信度来提升人类在协作决策任务中的表现?
- RQ2基于实证数据推导出的人类行为模型,如何指导为人类使用而优化的AI置信度变换?
- RQ3经修改的AI建议带来的性能提升是否在不同数据模态(图像、文本、表格)和用户群体中均成立?
- RQ4AI建议的准确率在多大程度上调节了置信度修改对人类表现的影响?
- RQ5所提方法是否可泛化至新用户群体,例如来自不同国家的参与者?
主要发现
- 在测试美国和英国参与者时,经修改的AI建议使人类准确率显著提高(p = 8.86 × 10⁻³)。
- 在两个群体中,人类对正确标签的置信度也显著提升(p = 4.43 × 10⁻¹¹)。
- 在AI建议本身已较准确的任务中(如艺术与城市识别,准确率分别为90%和87%),性能提升最为显著,表明建议质量与置信度修改之间存在协同效应。
- 使用AI建议的激活率平均提高了3.3%,表明用户参与度更高。
- 模拟结果与实证发现高度一致,验证了人类行为模型的预测能力。
- 结果表明,联合优化人机系统(而非孤立优化AI模型)可带来更优的实际效果。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。