[Paper Review] NIPS - Not Even Wrong? A Systematic Review of Empirically Complete Demonstrations of Algorithmic Effectiveness in the Machine Learning and Artificial Intelligence Literature
This systematic review evaluates the empirical rigor of algorithmic effectiveness claims in 121 supervised learning papers from the 2017 NeurIPS conference. It finds that only 2 out of 121 papers (1.6%) presented complete argumentative chains to support claims of superiority over state-of-the-art baselines, revealing widespread deficiencies in reporting standards and empirical validation in top-tier ML/AI research.
Objective: To determine the completeness of argumentative steps necessary to conclude effectiveness of an algorithm in a sample of current ML/AI supervised learning literature. Data Sources: Papers published in the Neural Information Processing Systems (NeurIPS, née NIPS) journal where the official record showed a 2017 year of publication. Eligibility Criteria: Studies reporting a (semi-)supervised model, or pre-processing fused with (semi-)supervised models for tabular data. Study Appraisal: Three reviewers applied the assessment criteria to determine argumentative completeness. The criteria were split into three groups, including: experiments (e.g real and/or synthetic data), baselines (e.g uninformed and/or state-of-art) and quantitative comparison (e.g. performance quantifiers with confidence intervals and formal comparison of the algorithm against baselines). Results: Of the 121 eligible manuscripts (from the sample of 679 abstracts), 99\% used real-world data and 29\% used synthetic data. 91\% of manuscripts did not report an uninformed baseline and 55\% reported a state-of-art baseline. 32\% reported confidence intervals for performance but none provided references or exposition for how these were calculated. 3\% reported formal comparisons. Limitations: The use of one journal as the primary information source may not be representative of all ML/AI literature. However, the NeurIPS conference is recognised to be amongst the top tier concerning ML/AI studies, so it is reasonable to consider its corpus to be representative of high-quality research. Conclusion: Using the 2017 sample of the NeurIPS supervised learning corpus as an indicator for the quality and trustworthiness of current ML/AI research, it appears that complete argumentative chains in demonstrations of algorithmic effectiveness are rare.
Motivation & Objective
- To assess the completeness of argumentative chains required to conclude that a new ML/AI algorithm is effective.
- To evaluate whether supervised learning papers in top-tier ML conferences provide sufficient empirical evidence to support claims of algorithmic superiority.
- To identify missing components in the reporting of baselines, performance metrics, and statistical comparisons in ML/AI research.
- To examine whether current publication practices in ML/AI encourage incomplete or misleading claims of effectiveness.
- To advocate for improved reporting standards to enhance the trustworthiness and reproducibility of ML/AI research.
Proposed method
- Systematically screened 679 abstracts from NeurIPS 2017 to identify eligible supervised learning studies.
- Selected 121 papers that reported (semi-)supervised models or pre-processing fused with such models for tabular data.
- Applied a three-tiered assessment framework: experiments (real/synthetic data), baselines (uninformed/state-of-the-art), and quantitative comparison (performance metrics with confidence intervals and formal testing).
- Three independent reviewers assessed each paper for argumentative completeness using predefined criteria.
- Used consensus and third-party arbitration to resolve disagreements during screening and full-text assessment.
- Conducted post-review statistical analysis to quantify reporting deficiencies across the corpus.
Experimental results
Research questions
- RQ1To what extent do NeurIPS 2017 supervised learning papers provide complete empirical arguments for algorithmic effectiveness?
- RQ2How frequently are uninformed baselines, state-of-the-art baselines, confidence intervals, and formal statistical comparisons reported in top-tier ML papers?
- RQ3What proportion of papers in the NeurIPS 2017 corpus present a fully testable and scientifically valid argument for algorithmic superiority?
- RQ4Why do most ML/AI papers fail to present complete argumentative chains despite the availability of established evaluation standards?
- RQ5What systemic issues in publication and review practices contribute to the prevalence of 'not even wrong' claims in ML/AI research?
Key findings
- Of 121 eligible papers, only 2 (1.6%) presented a complete argumentative chain supporting the claim that a new algorithm is effective.
- 99% of papers used real-world data, while only 29% used synthetic data, indicating a strong reliance on real data without systematic ablation.
- 91% of papers did not report an uninformed baseline, and 55% reported a state-of-the-art baseline, suggesting weak comparative rigor.
- Only 32% reported confidence intervals for performance metrics, and none provided justification or references for their calculation.
- Just 3% of papers conducted formal statistical comparisons to validate performance differences against baselines.
- The study concludes that complete empirical argumentation for algorithmic effectiveness is rare in high-impact ML/AI research, with most claims being 'not even wrong' due to missing foundational evidence.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.