[Paper Review] Data Pricing in Machine Learning Pipelines
This paper surveys data pricing principles in machine learning pipelines, focusing on three key stages: training data collection, collaborative model training, and model deployment. It proposes principled pricing mechanisms for raw data, labeled data, and model contributions using game-theoretic approaches like Shapley value, while identifying challenges in end-to-end revenue allocation, robustness to data replication, query-based pricing in competitive markets, and rigorous evaluation frameworks.
Machine learning is disruptive. At the same time, machine learning can only succeed by collaboration among many parties in multiple steps naturally as pipelines in an eco-system, such as collecting data for possible machine learning applications, collaboratively training models by multiple parties and delivering machine learning services to end users. Data is critical and penetrating in the whole machine learning pipelines. As machine learning pipelines involve many parties and, in order to be successful, have to form a constructive and dynamic eco-system, marketplaces and data pricing are fundamental in connecting and facilitating those many parties. In this article, we survey the principles and the latest research development of data pricing in machine learning pipelines. We start with a brief review of data marketplaces and pricing desiderata. Then, we focus on pricing in three important steps in machine learning pipelines. To understand pricing in the step of training data collection, we review pricing raw data sets and data labels. We also investigate pricing in the step of collaborative training of machine learning models, and overview pricing machine learning models for end users in the step of machine learning deployment. We also discuss a series of possible future directions.
Motivation & Objective
- To establish principled data pricing mechanisms for machine learning pipelines involving multiple collaborating parties.
- To address the unique challenges of pricing data as a non-rival, combinatorial, and heterogeneous asset with variable value across buyers.
- To survey and analyze pricing methods for raw data sets, data labeling, collaborative model contributions, and deployed machine learning models.
- To identify critical research gaps in end-to-end revenue allocation, axiomatic foundations of valuation, fine-grained data procurement, and evaluation of pricing models.
- To advocate for systematic frameworks that support dynamic, fair, and robust data and model exchange in ML ecosystems.
Proposed method
- Surveying existing research on data marketplaces and pricing desiderata, with a focus on three pipeline stages: data collection, collaborative training, and deployment.
- Applying cooperative game theory, particularly Shapley value, to quantify and price contributions in collaborative model training.
- Evaluating pricing models based on axioms such as symmetry, additivity, and zero element, and comparing alternatives like normalized Banzhaf value.
- Proposing query-based pricing models for selective data acquisition in multi-seller environments to enhance data diversity and budget efficiency.
- Recommending the development of simulation platforms to evaluate pricing models under realistic market behaviors, including adversarial and coalition strategies.
- Integrating principles of truthfulness, revenue maximization, and arbitrage-freeness into pricing model design.

Experimental results
Research questions
- RQ1How can data be priced in a way that fairly reflects its contribution to machine learning model performance across diverse buyers?
- RQ2What game-theoretic mechanisms can ensure fair and robust revenue allocation among contributors in collaborative model training?
- RQ3How can pricing models support fine-grained, query-based procurement of data in competitive marketplaces with overlapping data sets?
- RQ4What are the limitations of current pricing models when real-world behaviors—such as collusion or strategic bidding—disrupt idealized assumptions?
- RQ5How can end-to-end revenue allocation be systematically designed across the full machine learning pipeline, from data collection to model deployment?
Key findings
- Shapley value is widely used for contribution-based pricing in collaborative training but may not be optimal in all market contexts due to its reliance on the additivity axiom.
- Normalized Banzhaf value offers greater robustness to data replication attacks compared to Shapley value, suggesting context-dependent suitability of valuation methods.
- Existing pricing models often assume rational, non-coalitional behavior, but real-world market dynamics such as adversarial or coalition behaviors are under-evaluated in current research.
- There is a lack of systematic frameworks for end-to-end revenue allocation that account for contributions across all stages of the machine learning pipeline.
- Query-based pricing models exist but are currently limited to monopoly settings and need extension to multi-seller, competitive data marketplaces.
- Rigorous evaluation of pricing models remains underdeveloped, with most studies relying on oversimplified assumptions; simulation platforms are needed for realistic performance assessment.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.