[论文解读] Measurement as governance in and for responsible AI
本文主张,测量是人工智能系统中一种隐蔽的治理形式,社会、文化和政治价值观通过将公平性、鲁棒性和责任性等抽象概念具体化而被编码其中。通过应用建构效度——特别是内容效度和后果效度——本文揭示了理想建构与其测量之间不匹配所导致的与公平性相关的伤害,提出了一套框架,使人工智能治理更加透明和可问责。
Measurement of social phenomena is everywhere, unavoidably, in sociotechnical systems. This is not (only) an academic point: Fairness-related harms emerge when there is a mismatch in the measurement process between the thing we purport to be measuring and the thing we actually measure. However, the measurement process -- where social, cultural, and political values are implicitly encoded in sociotechnical systems -- is almost always obscured. Furthermore, this obscured process is where important governance decisions are encoded: governance about which systems are fair, which individuals belong in which categories, and so on. We can then use the language of measurement, and the tools of construct validity and reliability, to uncover hidden governance decisions. In particular, we highlight two types of construct validity, content validity and consequential validity, that are useful to elicit and characterize the feedback loops between the measurement, social construction, and enforcement of social categories. We then explore the constructs of fairness, robustness, and responsibility in the context of governance in and for responsible AI. Together, these perspectives help us unpack how measurement acts as a hidden governance process in sociotechnical systems. Understanding measurement as governance supports a richer understanding of the governance processes already happening in AI -- responsible or otherwise -- revealing paths to more effective interventions.
研究动机与目标
- 揭示人工智能系统中测量过程如何隐性地编码关于公平性、分类和责任的治理决策。
- 解决理论建构与其操作化测量之间不匹配所引发的与公平性相关的伤害问题。
- 证明测量并非中立,而是社会技术系统中嵌入价值观的治理场域。
- 提出一种基于建构效度(内容效度与后果效度)的框架,使这些隐藏的治理决策得以显现并可分析。
- 通过将社会科学视角融入技术治理,支持更有效、可问责和负责任的人工智能发展。
提出的方法
- 采用社会科学中的测量框架,分析不可观测建构(如“员工质量”、“毒性”)如何被操作化为可测量的代理指标。
- 应用两种建构效度:内容效度(测量是否反映预期建构?)与后果效度(测量引发何种社会后果?)
- 使用案例研究(如以过去薪资作为员工质量的代理指标)说明测量选择如何复制系统性不公。
- 绘制测量、社会建构类别与社会等级制度执行之间的反馈回路。
- 整合批判性种族理论、残障研究与组织治理的洞见,以语境化测量的影响。
- 将测量视为一种透镜,以揭示并批判算法系统中嵌入的治理决策。
实验结果
研究问题
- RQ1人工智能系统中的测量过程如何编码关于公平性、分类和责任的治理决策?
- RQ2理论建构与其操作化测量之间的不匹配以何种方式导致与公平性相关的伤害?
- RQ3如何运用内容效度与后果效度来诊断并改进人工智能系统中嵌入的治理?
- RQ4测量在算法系统中对身份类别(如种族、性别、残障)的社会建构与执行中扮演何种角色?
- RQ5基于测量的框架如何支持更透明和可问责的责任人工智能治理?
主要发现
- 测量并非中立的技术过程,而是社会、文化与政治价值观嵌入人工智能系统的根本治理场域。
- 与公平性相关的伤害通常并非源于算法偏见本身,而是源于理想建构(如“员工质量”)与其操作化(如“过去薪资”)之间的不匹配。
- 内容效度有助于评估测量模型是否充分代表其声称测量的理论建构。
- 后果效度揭示了测量如何塑造并强化社会类别与权力结构,包括通过行政暴力或符号暴力。
- 通过建构效度显式分析测量过程,可揭示隐藏的治理决策,从而实现更可问责的人工智能开发。
- 将测量重新定义为治理,有助于更深入地整合社会科学洞见至负责任的人工智能,从而提升透明度与伦理可问责性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。