Skip to main content
QUICK REVIEW

[Paper Review] Experimentation, Biased Learning, and Conjectural Variations in Competitive Dynamic Pricing

Bar Light, Wenyu Wang|arXiv (Cornell University)|Feb 13, 2026
Auction Theory and Applications0 citations
TL;DR

The paper analyzes multi-seller dynamic pricing with bandit feedback where sellers use two-point experiments and learn from own data. It shows that correlated experimentation induces biased demand learning, which endogenously selects a Conjectural Variations equilibrium, often supra-competitive, while independent experimentation leads to Nash equilibrium.

ABSTRACT

We study competitive dynamic pricing among multiple sellers, motivated by the rise of large-scale experimentation and algorithmic pricing in retail and online marketplaces. Sellers repeatedly set prices using simple learning rules and observe only their own prices and realized demand, even though demand depends on all sellers' prices and is subject to random shocks. Each seller runs two-point A/B price experiments, in the spirit of switchback-style designs, and updates a baseline price using a linear demand estimate fitted to its own data. Under certain conditions on demand, the resulting dynamics converge to a Conjectural Variations (CV) equilibrium, a classic static equilibrium notion in which each seller best responds under a conjecture that rivals' prices respond systematically to changes in its own price. Unlike standard CV models that treat conjectures as behavioral primitives, we show that these conjectures arise endogenously from the bias in demand learning induced by correlated experimentation (e.g., due to synchronized repricing schedules). This learning bias selects the long-run equilibrium, often leading to supra-competitive prices. Notably, we show that under independent experimentation, this bias vanishes and the learning dynamics converge to the standard Nash equilibrium. We provide simple sufficient conditions on demand for convergence in standard models and establish a finite-sample guarantee: up to logarithmic factors, the squared price error decays on the order of $T^{-1/2}$. Our results imply that in competitive markets, experimentation design can serve as a market design lever, selecting the equilibrium reached by practical learning algorithms.

Motivation & Objective

  • Motivate the study of competitive dynamic pricing with large-scale experimentation and bandit feedback in multiseller markets.
  • Characterize how self-learning price updates from two-point experiments converge to a Conjectural Variations equilibrium under correlated experimentation.
  • Identify conditions on demand under which learning dynamics converge and provide finite-sample convergence guarantees.
  • Show how correlated experimentation acts as a market design lever for equilibrium selection between CV and Nash outcomes.

Proposed method

  • Model a repeated pricing game with n sellers and bandit feedback where each seller observes only her own price and realized demand.
  • Introduce two-point price experimentation (baseline vs baseline plus small perturbation) and a linear regression demand estimator based on own data.
  • Propose Switchback Linear Demand Learning (SLDL): data collection via batch-wise randomized perturbations, OLS demand estimation, and partial-price updates toward estimated revenue-maximizing targets.
  • Define conjectural variations (CV) equilibrium via a conjecture matrix A capturing rivals’ assumed cross-price responses, and derive first-order conditions for CV equilibria.
  • Show that correlated experimentation induces a bias in demand estimation that mimics co-movement of rivals’ prices, leading to CV equilibrium as the learning outcome.
  • Establish finite-sample guarantees: mean squared price error decays at rate T^{-1/2} up to logarithmic factors under stability conditions.

Experimental results

Research questions

  • RQ1Can a simple bandit-feedback pricing algorithm in a multi-seller setting converge to a CV equilibrium?
  • RQ2How does the correlation structure of price experimentation affect equilibrium selection and pricing outcomes?
  • RQ3Under what conditions on demand do learning dynamics converge, and what is the rate of convergence?
  • RQ4What is the endogenous mechanism by which learning bias from correlated experimentation generates CV conjectures?
  • RQ5How does independent (uncorrelated) experimentation affect convergence to Nash equilibrium?

Key findings

  • The learning dynamics converge to a CV equilibrium when sellers follow the proposed two-point experimentation and linear-demand learning under certain demand conditions.
  • The limiting conjecture matrix is endogenously determined by the statistical structure of correlated experimentation, not imposed as an external primitive.
  • When experimentation is uncorrelated across sellers, the bias vanishes and dynamics converge to the standard Nash equilibrium.
  • The paper provides simple sufficient conditions on demand (via derivatives) for convergence in standard models, with explicit stability bounds for linear and multinomial logit models.
  • Under the stated stability criteria, there is a finite-sample convergence guarantee: the mean squared price error decays on the order of T^{-1/2} up to logarithmic factors.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.