Skip to main content
QUICK REVIEW

[论文解读] Ethical behavior in humans and machines -- Evaluating training data quality for beneficial machine learning

Thilo Hagendorff|arXiv (Cornell University)|Aug 26, 2020
Ethics and Social Impacts of AI参考文献 99被引用 4
一句话总结

本文提出了一种用于评估监督机器学习中训练数据质量的伦理框架,认为数据质量必须超越技术指标,涵盖伦理维度。通过从社会和心理学角度分析行为数据,本文提出一种选择性数据过滤机制,以替代‘n = 全部’的大数据范式,从而实现更负责任且有益的人工智能系统。

ABSTRACT

Machine behavior that is based on learning algorithms can be significantly influenced by the exposure to data of different qualities. Up to now, those qualities are solely measured in technical terms, but not in ethical ones, despite the significant role of training and annotation data in supervised machine learning. This is the first study to fill this gap by describing new dimensions of data quality for supervised machine learning applications. Based on the rationale that different social and psychological backgrounds of individuals correlate in practice with different modes of human-computer-interaction, the paper describes from an ethical perspective how varying qualities of behavioral data that individuals leave behind while using digital technologies have socially relevant ramification for the development of machine learning applications. The specific objective of this study is to describe how training data can be selected according to ethical assessments of the behavior it originates from, establishing an innovative filter regime to transition from the big data rationale n = all to a more selective way of processing data for training sets in machine learning. The overarching aim of this research is to promote methods for achieving beneficial machine learning applications that could be widely useful for industry as well as academia.

研究动机与目标

  • 解决监督机器学习中训练数据质量缺乏伦理评估的问题。
  • 探讨不同的社会和心理背景如何影响人机交互及数据生成。
  • 提出从‘n = 全部’的大数据原则向更具选择性、伦理导向的数据整理过程转变。
  • 通过经过伦理评估的训练数据,为开发有益的机器学习应用奠定基础。
  • 提出一种新颖的过滤机制,基于其所代表行为的伦理质量来评估数据。

提出的方法

  • 以数字交互过程中生成的行为数据为基础,进行伦理数据评估。
  • 引入新的数据质量维度,聚焦于伦理影响,而非仅技术性能。
  • 提出一种基于数据生产者个体行为伦理性的过滤机制。
  • 借鉴社会与心理学理论,从伦理角度关联个体背景与数据质量。
  • 将机器学习中的数据选择重新定义为一种道德与社会责任,而不仅是技术优化。
  • 以经过伦理审查的选择性数据整理策略,取代默认的‘全部数据’训练方法。

实验结果

研究问题

  • RQ1个体之间的社会与心理差异在多大程度上影响数字系统中行为数据的伦理质量?
  • RQ2在监督机器学习的训练数据中,可以识别并衡量哪些伦理维度的数据质量?
  • RQ3对人类行为的伦理评估在多大程度上可指导机器学习模型训练数据的选择?
  • RQ4从‘n = 全部’的数据收集方式转向选择性、经伦理评估的数据,如何改善机器学习结果?
  • RQ5可以构建何种框架,以将伦理考量整合到有益人工智能的数据整理过程中?

主要发现

  • 本文指出,训练数据质量必须包含超越技术指标(如准确率或平衡性)的伦理维度。
  • 研究证明,个体行为数据反映了其社会与心理背景,从而影响数据使用的伦理后果。
  • 本研究建立了基于其所代表行为伦理质量的数据评估概念框架。
  • 提出一种新颖的过滤机制,摆脱‘n = 全部’的数据收集模式,转向更具选择性、伦理导向的数据筛选。
  • 该研究为通过伦理评估的训练数据开发有益的机器学习系统提供了基础方法。
  • 该工作将伦理数据整理定位为负责任人工智能发展的关键,对产业界与学术界均具有深远影响。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。