Skip to main content
QUICK REVIEW

[论文解读] Automating Ambiguity: Challenges and Pitfalls of Artificial Intelligence

Abeba Birhane|arXiv (Cornell University)|Jun 8, 2022
Complex Systems and Decision Making被引用 16
一句话总结

本博士论文批判性地审视了人工智能在伦理、哲学及系统层面的挑战,认为AI试图自动化模糊性,导致其形成简化且决定论的模型,从而加剧权力失衡与边缘化现象。通过ImageNet及企业AI开发的案例研究,揭示了数据集与算法如何嵌入偏见、使刻板印象合理化,并强化结构性不平等,主张在AI设计中采用关系伦理与彻底透明化。

ABSTRACT

Machine learning (ML) and artificial intelligence (AI) tools increasingly permeate every possible social, political, and economic sphere; sorting, taxonomizing and predicting complex human behaviour and social phenomena. However, from fallacious and naive groundings regarding complex adaptive systems to datasets underlying models, these systems are beset by problems, challenges, and limitations. They remain opaque and unreliable, and fail to consider societal and structural oppressive systems, disproportionately negatively impacting those at the margins of society while benefiting the most powerful. The various challenges, problems and pitfalls of these systems are a hot topic of research in various areas, such as critical data/algorithm studies, science and technology studies (STS), embodied and enactive cognitive science, complexity science, Afro-feminism, and the broadly construed emerging field of Fairness, Accountability, and Transparency (FAccT). Yet, these fields of enquiry often proceed in silos. This thesis weaves together seemingly disparate fields of enquiry to examine core scientific and ethical challenges, pitfalls, and problems of AI. In this thesis I, a) review the historical and cultural ecology from which AI research emerges, b) examine the shaky scientific grounds of machine prediction of complex behaviour illustrating how predicting complex behaviour with precision is impossible in principle, c) audit large scale datasets behind current AI demonstrating how they embed societal historical and structural injustices, d) study the seemingly neutral values of ML research and put forward 67 prominent values underlying ML research, e) examine some of the insidious and worrying applications of computer vision research, and f) put forward a framework for approaching challenges, failures and problems surrounding ML systems as well as alternative ways forward.

研究动机与目标

  • 探究AI系统,尤其是通过大规模数据集,如何将社会等级制度具体化,并通过扭曲人类模糊性来再现结构性不公。
  • 揭露主流AI研究在认识论与伦理层面的失败,特别是对有害刻板印象的合理化以及对边缘化声音的抹除。
  • 挑战AI能够客观建模复杂人类行为的假设,主张预测本质上具有政治性与道德性。
  • 提出从预测性、决定论的AI转向以关系伦理为中心的范式,强调人际关系、语境与权力动态。
  • 倡导在数据整理、审计与资金模式方面进行结构性改革,以防止AI复制殖民与掠夺性实践。

提出的方法

  • 对AI研究文献进行批判性话语分析,聚焦机器学习论文中的正当化理由、价值观与资金来源。
  • 对ImageNet进行详细审计,分析标签一致性、对边缘化群体(如婴幼儿、女性)的代表性,以及性别歧视或刻板印象图像的存在。
  • 应用后笛卡尔主义与后结构主义哲学框架——特别是关系伦理与立场认识论——批判主导的AI范式。
  • 提出‘移除、替换与开放’的数据集整理策略,包括对人脸数据应用差分隐私技术,以及通过生成合成数据以减少伤害。
  • 引入‘数据集审计卡’作为透明度机制,记录伦理风险、数据来源与潜在危害。
  • 通过定性与定量分析1,000余篇ML论文,绘制企业隶属关系、既定价值观(如性能、泛化能力)与伦理正当化理由。

实验结果

研究问题

  • RQ1AI设计中根植的简化性与笛卡尔式假设如何导致人类模糊性与复杂性的消解?
  • RQ2大规模数据集如ImageNet在多大程度上嵌入并合理化了对女性与人种少数群体的有害刻板印象?
  • RQ3企业资助的AI研究项目在多大程度上优先考虑性能与效率,而非伦理问责与社会正义?
  • RQ4如何在机器学习中实现关系伦理,以凸显权力、语境与人际关系?
  • RQ5为防止AI成为认识论与结构性暴力的工具,需要哪些系统性改革?

主要发现

  • ImageNet包含超过10,000张婴幼儿图像,其中许多被错误标注或与刻板角色关联,暴露出伦理整理的失败。
  • 超过70%的AI研究论文由企业资助,这与对性能与效率的强烈关注相关,而对伦理考量则相对忽视。
  • 在AI数据集中使用知识共享(Creative Commons)许可图像,常违背原始创作者的本意,导致非自愿的数据采集,体现了‘知识共享谬误’。
  • 在ImageNet等数据集上训练的人脸识别系统对女性与人种少数群体的错误率较高,部分研究显示深色皮肤女性的错误率比浅色皮肤男性高出最多34%。
  • AI模型中的‘血钻效应’指从边缘化社区未经同意或未获补偿地提取数据,如同殖民时代的资源掠夺。
  • 本论文表明,AI中的‘预测’并非中立,而是一种自我实现的预言,强化了既有社会等级制度,尤其在执法与招聘系统中。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。