Skip to main content
QUICK REVIEW

[Paper Review] NIPS - Not Even Wrong? A Systematic Review of Empirically Complete Demonstrations of Algorithmic Effectiveness in the Machine Learning and Artificial Intelligence Literature

Franz J. Király, Bilal A. Mateen|arXiv (Cornell University)|Dec 18, 2018
Explainable Artificial Intelligence (XAI)76 references5 citations
TL;DR

This systematic review evaluates the empirical rigor of algorithmic effectiveness claims in 121 supervised learning papers from the 2017 NeurIPS conference. It finds that only 2 out of 121 papers (1.6%) presented complete argumentative chains to support claims of superiority over state-of-the-art baselines, revealing widespread deficiencies in reporting standards and empirical validation in top-tier ML/AI research.

ABSTRACT

Objective: To determine the completeness of argumentative steps necessary to conclude effectiveness of an algorithm in a sample of current ML/AI supervised learning literature. Data Sources: Papers published in the Neural Information Processing Systems (NeurIPS, née NIPS) journal where the official record showed a 2017 year of publication. Eligibility Criteria: Studies reporting a (semi-)supervised model, or pre-processing fused with (semi-)supervised models for tabular data. Study Appraisal: Three reviewers applied the assessment criteria to determine argumentative completeness. The criteria were split into three groups, including: experiments (e.g real and/or synthetic data), baselines (e.g uninformed and/or state-of-art) and quantitative comparison (e.g. performance quantifiers with confidence intervals and formal comparison of the algorithm against baselines). Results: Of the 121 eligible manuscripts (from the sample of 679 abstracts), 99\% used real-world data and 29\% used synthetic data. 91\% of manuscripts did not report an uninformed baseline and 55\% reported a state-of-art baseline. 32\% reported confidence intervals for performance but none provided references or exposition for how these were calculated. 3\% reported formal comparisons. Limitations: The use of one journal as the primary information source may not be representative of all ML/AI literature. However, the NeurIPS conference is recognised to be amongst the top tier concerning ML/AI studies, so it is reasonable to consider its corpus to be representative of high-quality research. Conclusion: Using the 2017 sample of the NeurIPS supervised learning corpus as an indicator for the quality and trustworthiness of current ML/AI research, it appears that complete argumentative chains in demonstrations of algorithmic effectiveness are rare.

Motivation & Objective

  • To assess the completeness of argumentative chains required to conclude that a new ML/AI algorithm is effective.
  • To evaluate whether supervised learning papers in top-tier ML conferences provide sufficient empirical evidence to support claims of algorithmic superiority.
  • To identify missing components in the reporting of baselines, performance metrics, and statistical comparisons in ML/AI research.
  • To examine whether current publication practices in ML/AI encourage incomplete or misleading claims of effectiveness.
  • To advocate for improved reporting standards to enhance the trustworthiness and reproducibility of ML/AI research.

Proposed method

  • Systematically screened 679 abstracts from NeurIPS 2017 to identify eligible supervised learning studies.
  • Selected 121 papers that reported (semi-)supervised models or pre-processing fused with such models for tabular data.
  • Applied a three-tiered assessment framework: experiments (real/synthetic data), baselines (uninformed/state-of-the-art), and quantitative comparison (performance metrics with confidence intervals and formal testing).
  • Three independent reviewers assessed each paper for argumentative completeness using predefined criteria.
  • Used consensus and third-party arbitration to resolve disagreements during screening and full-text assessment.
  • Conducted post-review statistical analysis to quantify reporting deficiencies across the corpus.

Experimental results

Research questions

  • RQ1To what extent do NeurIPS 2017 supervised learning papers provide complete empirical arguments for algorithmic effectiveness?
  • RQ2How frequently are uninformed baselines, state-of-the-art baselines, confidence intervals, and formal statistical comparisons reported in top-tier ML papers?
  • RQ3What proportion of papers in the NeurIPS 2017 corpus present a fully testable and scientifically valid argument for algorithmic superiority?
  • RQ4Why do most ML/AI papers fail to present complete argumentative chains despite the availability of established evaluation standards?
  • RQ5What systemic issues in publication and review practices contribute to the prevalence of 'not even wrong' claims in ML/AI research?

Key findings

  • Of 121 eligible papers, only 2 (1.6%) presented a complete argumentative chain supporting the claim that a new algorithm is effective.
  • 99% of papers used real-world data, while only 29% used synthetic data, indicating a strong reliance on real data without systematic ablation.
  • 91% of papers did not report an uninformed baseline, and 55% reported a state-of-the-art baseline, suggesting weak comparative rigor.
  • Only 32% reported confidence intervals for performance metrics, and none provided justification or references for their calculation.
  • Just 3% of papers conducted formal statistical comparisons to validate performance differences against baselines.
  • The study concludes that complete empirical argumentation for algorithmic effectiveness is rare in high-impact ML/AI research, with most claims being 'not even wrong' due to missing foundational evidence.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.