[论文解读] A Taxonomy of Human and ML Strengths in Decision-Making to Investigate Human-ML Complementarity
本文提出了一种人类与机器学习(ML)在决策中优势的分类法,以澄清人类-ML互补性的机制。通过分析不同的认知与统计能力,作者构建了一个最优聚合人类与ML判断的数学框架,表明当目标不同时,互补性最强,且最优协作方式是根据实例选择表现更优的代理,而非凸组合。
Hybrid human-ML systems increasingly make consequential decisions in a wide range of domains. These systems are often introduced with the expectation that the combined human-ML system will achieve complementary performance, that is, the combined decision-making system will be an improvement compared with either decision-making agent in isolation. However, empirical results have been mixed, and existing research rarely articulates the sources and mechanisms by which complementary performance is expected to arise. Our goal in this work is to provide conceptual tools to advance the way researchers reason and communicate about human-ML complementarity. Drawing upon prior literature in human psychology, machine learning, and human-computer interaction, we propose a taxonomy characterizing distinct ways in which human and ML-based decision-making can differ. In doing so, we conceptually map potential mechanisms by which combining human and ML decision-making may yield complementary performance, developing a language for the research community to reason about design of hybrid systems in any decision-making domain. To illustrate how our taxonomy can be used to investigate complementarity, we provide a mathematical aggregation framework to examine enabling conditions for complementarity. Through synthetic simulations, we demonstrate how this framework can be used to explore specific aspects of our taxonomy and shed light on the optimal mechanisms for combining human-ML judgments
研究动机与目标
- 为解决人类-ML系统在某些情况下优于单一代理却缺乏概念清晰性的原因。
- 识别并分类人类与ML模型在决策情境中的独特优势与劣势。
- 开发一个正式框架,用于分析人类-ML协作在何时以及如何实现互补性能。
- 通过提供共享语言与分析工具,指导研究人员和实践者设计混合系统,以预测有效协作。
提出的方法
- 基于心理学、机器学习与人机交互的洞见,提出人类与ML决策优势的分类法。
- 将互补性分类为跨实例与单实例类型,区分协作提升性能的情境。
- 开发凸组合聚合框架,将联合决策建模为人类与ML预测的加权平均。
- 引入性能度量 $c_{\text{across}}$ 与 $c_{\text{within}}$,以量化在不同目标对齐条件下的互补性。
- 使用合成模拟探索互补性出现的条件,变化目标对齐程度(以 $b$ 参数化)。
- 通过比较联合性能与个体代理表现,分析最优聚合机制,发现当目标不同时,基于委托的策略优于凸组合。
实验结果
研究问题
- RQ1人类与ML决策在优势与劣势方面存在哪些不同的方式?
- RQ2在何种条件下,人类-ML协作能产生互补性能,而非冗余或性能下降?
- RQ3人类与ML目标的对齐程度如何影响互补性的潜力?
- RQ4聚合人类与ML判断以最大化性能的最优机制是什么?
- RQ5是否存在一个统一框架,能够解释并预测在多样化决策领域中有效的人类-ML协作?
主要发现
- 当人类与ML目标完全对齐时($b=1$),互补性能最低,因为缺乏协作激励。
- 当目标不同时($b \neq 1$),跨实例互补性($c_{\text{across}}$)相对较高,表明团队性能增益潜力强。
- 在目标错位下,单实例互补性($c_{\text{within}}$)仍保持较低水平,表明最优协作并不要求双方在每个实例中都参与。
- 最优聚合机制并非判断的凸组合,而是基于实例将决策委托给表现更优的代理。
- 合成模拟证实,当人类与ML模型目标相异时,基于委托的协作优于固定权重平均。
- 该框架能够生成关于最优人类-ML协作的假设,并量化性能增益与实施成本之间的权衡。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。