Skip to main content
QUICK REVIEW

[Paper Review] Expanding search in the space of empirical ML

Bronwyn Woods|arXiv (Cornell University)|Dec 4, 2018
Reinforcement Learning in Robotics9 references4 citations
TL;DR

This paper argues for expanding machine learning conference culture to value empirical research beyond algorithmic novelty, advocating for systematic recognition of data, target, and experimental synthesis work. It proposes redefining novelty, creating dedicated tracks for synthesis and open-world experimentation, and incentivizing industry-academic collaboration to improve reproducibility and real-world applicability of ML research.

ABSTRACT

As researchers and practitioners of applied machine learning, we are given a set of requirements on the problem to be solved, the plausibly obtainable data, and the computational resources available. We aim to find (within those bounds) reliably useful combinations of problem, data, and algorithm. An emphasis on algorithmic or technical novelty in ML conference publications leads to exploration of one dimension of this space. Data collection and ML deployment at scale in industry settings offers an environment for exploring the others. Our conferences and reviewing criteria can better support empirical ML by soliciting and incentivizing experimentation and synthesis independent of algorithmic innovation.

Motivation & Objective

  • Address the imbalance in ML conference incentives that prioritize algorithmic novelty over empirical exploration of data and target dimensions.
  • Highlight the underappreciated role of data and target variability in shaping real-world ML performance, beyond benchmark-driven algorithmic comparisons.
  • Propose structural changes to conference culture and reviewing criteria to support and incentivize non-algorithmic empirical contributions such as synthesis, observation, and open-world experimentation.
  • Strengthen reproducibility and practical relevance in ML by recognizing the value of large-scale, real-world data and system performance analysis from industry practitioners.
  • Bridge the gap between academic innovation and industrial deployment by creating pathways for industry researchers to contribute empirical insights without requiring algorithmic novelty.

Proposed method

  • Reframe the ML research space as a three-dimensional search space: target, data, and algorithm, arguing that all dimensions deserve equal methodological attention.
  • Propose redefining 'novelty' in conference submissions to include insights from comparative empirical studies of existing algorithms across diverse data and targets.
  • Advocate for dedicated conference tracks—modeled after the IEEE S&P SoK track and Distill’s publication model—focused on systematic synthesis, taxonomies, and evidence-based challenges to long-held beliefs.
  • Encourage solicitation of open-world experimental results from industry researchers who have access to real-world data and system performance under diverse conditions.
  • Support the creation of shared, representative datasets and improved tooling to reduce barriers to reproducible empirical work, while emphasizing that such tools do not replace the need for systematic experimentation.
  • Promote collaboration between algorithmic researchers and open-world experimenters to identify and address practical gaps in existing models.

Experimental results

Research questions

  • RQ1How can ML conferences better recognize and incentivize empirical research that does not involve algorithmic innovation?
  • RQ2What role do data and target dimensions play in shaping the performance and generalizability of ML systems, and why are they underexplored in current research culture?
  • RQ3How can existing conference formats and reviewing criteria be adapted to support synthesis and observational studies of known algorithms across diverse datasets and real-world settings?
  • RQ4What structural changes are needed to integrate insights from industry practitioners into academic discourse, especially when their work lacks algorithmic novelty?
  • RQ5In what ways can open-world experimentation and shared datasets improve the reproducibility and practical relevance of ML research?

Key findings

  • Empirical ML research is systematically undervalued in current conference culture, which prioritizes algorithmic novelty over insights from data, target, and system-level experimentation.
  • Benchmark datasets are often treated as fixed and representative, but they introduce sampling variability and bias that significantly affect performance measurements and generalizability.
  • Reproducibility is undermined when experimental variability—such as hyperparameter tuning, random seeds, and codebase differences—is not systematically reported or controlled.
  • Existing work shows that with sufficient tuning, older models can outperform newer ones, indicating that reported gains may stem from experimental noise rather than algorithmic advances.
  • Synthesis and observational studies—such as those in the IEEE S&P SoK track or Distill—demonstrate that non-algorithmic contributions can yield high-impact insights into system behavior, performance patterns, and long-standing assumptions.
  • Industry researchers possess critical empirical knowledge from real-world deployments, but their contributions are often excluded from academic discourse due to misaligned incentives and publication norms.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.