[Paper Review] Ethical behavior in humans and machines -- Evaluating training data quality for beneficial machine learning
This paper introduces an ethical framework for evaluating training data quality in supervised machine learning, arguing that data quality must extend beyond technical metrics to include ethical dimensions. By analyzing behavioral data through social and psychological lenses, it proposes a selective data filtering regime to replace the 'n = all' big data paradigm, enabling more responsible and beneficial AI systems.
Machine behavior that is based on learning algorithms can be significantly influenced by the exposure to data of different qualities. Up to now, those qualities are solely measured in technical terms, but not in ethical ones, despite the significant role of training and annotation data in supervised machine learning. This is the first study to fill this gap by describing new dimensions of data quality for supervised machine learning applications. Based on the rationale that different social and psychological backgrounds of individuals correlate in practice with different modes of human-computer-interaction, the paper describes from an ethical perspective how varying qualities of behavioral data that individuals leave behind while using digital technologies have socially relevant ramification for the development of machine learning applications. The specific objective of this study is to describe how training data can be selected according to ethical assessments of the behavior it originates from, establishing an innovative filter regime to transition from the big data rationale n = all to a more selective way of processing data for training sets in machine learning. The overarching aim of this research is to promote methods for achieving beneficial machine learning applications that could be widely useful for industry as well as academia.
Motivation & Objective
- To address the lack of ethical evaluation in training data quality for supervised machine learning.
- To explore how diverse social and psychological backgrounds influence human-computer interaction and data generation.
- To propose a shift from the big data principle 'n = all' to a more selective, ethically informed data curation process.
- To establish a foundation for developing beneficial machine learning applications through ethically assessed training data.
- To provide a novel filter regime that evaluates data based on the ethical quality of the behavior it represents.
Proposed method
- Analyzes behavioral data generated during digital interactions as a basis for ethical data assessment.
- Introduces new dimensions of data quality focused on ethical implications rather than technical performance alone.
- Proposes a filter regime that evaluates data based on the ethical behavior of the individuals who produced it.
- Draws on social and psychological theories to correlate individual backgrounds with data quality from an ethical standpoint.
- Reframes data selection in machine learning as a moral and societal responsibility, not just a technical optimization.
- Replaces the default 'all-data' training approach with a selective, ethically vetted data curation strategy.
Experimental results
Research questions
- RQ1How do social and psychological differences among individuals influence the ethical quality of behavioral data in digital systems?
- RQ2What ethical dimensions of data quality can be identified and measured in supervised machine learning training data?
- RQ3To what extent can ethical assessments of human behavior inform the selection of training data for machine learning models?
- RQ4How can a shift from 'n = all' data collection to selective, ethically evaluated data improve machine learning outcomes?
- RQ5What framework can be developed to integrate ethical considerations into data curation for beneficial AI?
Key findings
- The paper identifies that training data quality must include ethical dimensions beyond technical metrics such as accuracy or balance.
- It demonstrates that individual behavioral data reflects social and psychological backgrounds that influence ethical implications of data use.
- The study establishes a conceptual framework for evaluating data based on the ethical quality of the behavior it represents.
- It proposes a novel filter regime that moves away from the 'n = all' data collection model toward selective, ethically informed data selection.
- The research contributes a foundational approach for developing beneficial machine learning systems through ethically assessed training data.
- The work positions ethical data curation as essential for responsible AI development, with implications for both industry and academia.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.