[Paper Review] Garbage In, Garbage Out? Do Machine Learning Application Papers in Social Computing Report Where Human-Labeled Training Data Comes From?
The paper audits ML classification papers on Twitter to see if they report how human-labeled training data were created, and finds substantial variation and often missing details about annotators, training, and data provenance.
Many machine learning projects for new application areas involve teams of humans who label data for a particular purpose, from hiring crowdworkers to the paper's authors labeling the data themselves. Such a task is quite similar to (or a form of) structured content analysis, which is a longstanding methodology in the social sciences and humanities, with many established best practices. In this paper, we investigate to what extent a sample of machine learning application papers in social computing --- specifically papers from ArXiv and traditional publications performing an ML classification task on Twitter data --- give specific details about whether such best practices were followed. Our team conducted multiple rounds of structured content analysis of each paper, making determinations such as: Does the paper report who the labelers were, what their qualifications were, whether they independently labeled the same items, whether inter-rater reliability metrics were disclosed, what level of training and/or instructions were given to labelers, whether compensation for crowdworkers is disclosed, and if the training data is publicly available. We find a wide divergence in whether such practices were followed and documented. Much of machine learning research and education focuses on what is done once a "gold standard" of training data is available, but we discuss issues around the equally-important aspect of whether such data is reliable in the first place.
Motivation & Objective
- Assess how ML application papers in social computing report the provenance of human-labeled training data.
- Evaluate transparency of annotator sources, qualifications, training, and compensation.
- Examine reporting of inter-annotator reliability and data availability.
- Highlight implications for data reliability and research integrity in supervised ML applications.
Proposed method
- Assemble a corpus of Twitter-classification ML papers from ArXiv and Scopus (approx. 494 ArXiv papers; 29 Scopus papers).
- Use a six-person labeling team to perform structured content analysis on each paper about data labeling practices.
- Apply a dual-round labeling process with reconciliation to determine reporting on annotators, training, definitions, and data availability.
- Develop raw and normalized information scores to quantify reporting of labeling details.
- Compute inter-rater reliability (IRR) as mean percent agreement across rounds (66.67% in round 1; 84.80% in round 2).
- Make datasets and code available on GitHub and Zenodo for reproducibility.
Experimental results
Research questions
- RQ1Do ML papers performing Twitter classification disclose whether the training data were labeled by humans?
- RQ2Who were the annotators (authors, crowdworkers, experts, etc.) and how were they recruited?
- RQ3What level of training, instructions, and inter-annotator reliability metrics are reported?
- RQ4Is crowdworker compensation disclosed and is the training data publicly accessible?
Key findings
- Most papers involved an original classification task (142 Yes, 17 No, 5 Unsure).
- Among papers with human annotation, 93 reported human annotation (Yes), 46 did not (No), 4 Unsure.
- For papers using original human annotation, 72 reported using original annotation (Yes), 21 did not (No), 3 Unsure.
- Annotator sources were diverse: authors themselves were the source in 22 papers (29.73%), and “no information” was common (24.32%); experts/professionals accounted for 21.62%, Amazon Mechanical Turk 4.05%, other crowdwork 10.81%, and other 9.46%.
- About half of papers using original human annotation specified the number of annotators (Yes 41; No 44.60% not specified).
- Formal instructions or definitions were reported in 32 papers (43.24%), while 35 (47.30%) provided no information on instructions; 7 (9.46%) stated no instructions beyond the question text.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.