[论文解读] Advancing Human-AI Complementarity: The Impact of User Expertise and Algorithmic Tuning on Joint Decision Making
本研究探讨了用户专业知识水平与算法调优如何影响人类-人工智能协作在血管标注任务中的表现。研究发现,当人工智能被调优以减少假阴性(与人类优势相匹配)时,对中等水平表现用户的性能提升最大;而专家用户则从个性化调优中获益最多,尽管有AI支持,新手用户的收益仍有限。
Human-AI collaboration for decision-making strives to achieve team performance that exceeds the performance of humans or AI alone. However, many factors can impact success of Human-AI teams, including a user's domain expertise, mental models of an AI system, trust in recommendations, and more. This work examines users' interaction with three simulated algorithmic models, all with similar accuracy but different tuning on their true positive and true negative rates. Our study examined user performance in a non-trivial blood vessel labeling task where participants indicated whether a given blood vessel was flowing or stalled. Our results show that while recommendations from an AI-Assistant can aid user decision making, factors such as users' baseline performance relative to the AI and complementary tuning of AI error types significantly impact overall team performance. Novice users improved, but not to the accuracy level of the AI. Highly proficient users were generally able to discern when they should follow the AI recommendation and typically maintained or improved their performance. Mid-performers, who had a similar level of accuracy to the AI, were most variable in terms of whether the AI recommendations helped or hurt their performance. In addition, we found that users' perception of the AI's performance relative on their own also had a significant impact on whether their accuracy improved when given AI recommendations. This work provides insights on the complexity of factors related to Human-AI collaboration and provides recommendations on how to develop human-centered AI algorithms to complement users in decision-making tasks.
研究动机与目标
- 理解用户专业知识水平如何影响人工智能辅助在联合决策任务中的有效性。
- 探究针对特定真正阳性率与真阴性率调优人工智能模型,如何影响人机协作团队的表现。
- 考察用户对人工智能性能相对于自身表现的感知,如何影响使用人工智能建议时的决策准确性。
- 探索心智模型与信任在塑造人机协作结果中的作用。
- 为设计能够互补人类优势并提升团队表现的人工智能助手提供设计指南。
提出的方法
- 在Stall Catchers公民科学平台上开展一项受控实验,共150次试验,模拟血管标注任务。
- 使用三个整体准确率相同但真正阳性率与真阴性率权衡不同的人工智能模型。
- 根据用户在无AI情况下的基线准确率,将用户按专业知识水平分为三类:新手、中等水平表现者与专家。
- 收集用户反馈与感知数据,以分析信任度、心智模型及决策行为。
- 分析用户决策与人工智能建议的一致性,以评估其对准确率与一致性的影响力。
- 使用聚类方法按表现水平对用户分组,并在各组内评估人工智能调优的影响。

实验结果
研究问题
- RQ1用户专业知识水平如何影响人工智能建议对决策准确率的影响?
- RQ2调优人工智能模型的真正阳性率与真阴性率,如何影响人机协作中的团队表现?
- RQ3用户对人工智能性能相对于自身表现的感知,在多大程度上影响其使用人工智能辅助时的决策行为?
- RQ4人类与人工智能错误模式之间的互补性如何影响整体团队准确率?
- RQ5是否可对人工智能调优进行优化,以增强特定用户专业知识群体的表现?
主要发现
- 基线准确率与人工智能相近的中等水平表现用户,其结果波动最大——在某些情况下人工智能辅助提升了表现,而在其他情况下反而降低了表现。
- 当人工智能被调优以减少假阴性(以增加假阳性为代价)时,用户准确率得以提升,因为他们能更轻松地拒绝错误的建议。
- 新手用户在人工智能辅助下表现有所提升,但未达到人工智能的准确率水平,表明当用户基线表现较低时,人工智能带来的收益有限。
- 高水平用户通过有选择性地采纳人工智能建议,维持或提升了自身表现,表明其具备较强的评估人工智能可靠性能力。
- 用户对人工智能性能相对于自身表现的感知,显著影响其采纳建议后是否提升准确率。
- 当人工智能调优与人类优势相匹配时,团队表现达到最大化——特别是当用户更擅长识别血流血管,而人工智能被调优以最小化假阴性时。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。