[논문 리뷰] Experimentation, Biased Learning, and Conjectural Variations in Competitive Dynamic Pricing
이 논문은 다중 판매자 환경에서 밴딧 피드백 하에 두 점 실험과 자체 데이터로 학습하는 동적 가격 결정 문제를 분석한다. 상관된 실험은 수요 학습에 편향을 만들어 Conjectural Variations 균형을 내생적으로 유도하는 반면, 독립적 실험은 Nash 균형으로 수렴한다.
We study competitive dynamic pricing among multiple sellers, motivated by the rise of large-scale experimentation and algorithmic pricing in retail and online marketplaces. Sellers repeatedly set prices using simple learning rules and observe only their own prices and realized demand, even though demand depends on all sellers' prices and is subject to random shocks. Each seller runs two-point A/B price experiments, in the spirit of switchback-style designs, and updates a baseline price using a linear demand estimate fitted to its own data. Under certain conditions on demand, the resulting dynamics converge to a Conjectural Variations (CV) equilibrium, a classic static equilibrium notion in which each seller best responds under a conjecture that rivals' prices respond systematically to changes in its own price. Unlike standard CV models that treat conjectures as behavioral primitives, we show that these conjectures arise endogenously from the bias in demand learning induced by correlated experimentation (e.g., due to synchronized repricing schedules). This learning bias selects the long-run equilibrium, often leading to supra-competitive prices. Notably, we show that under independent experimentation, this bias vanishes and the learning dynamics converge to the standard Nash equilibrium. We provide simple sufficient conditions on demand for convergence in standard models and establish a finite-sample guarantee: up to logarithmic factors, the squared price error decays on the order of $T^{-1/2}$. Our results imply that in competitive markets, experimentation design can serve as a market design lever, selecting the equilibrium reached by practical learning algorithms.
연구 동기 및 목표
- 다중 판매자 시장에서 대규모 실험과 밴딧 피드백을 활용한 경쟁적 동적 가격 결정 연구를 동기 부여한다.
- 상관된 실험 하에서 두 점 실험으로부터의 자체 학습 가격 업데이트가 어떻게 Conjectural Variations 균형으로 수렴하는지 특징지운다.
- 학습 역학이 수렴하는 수요 조건을 식별하고 유한 샘플 수렴 보장을 제시한다.
- 상관된 실험이 CV와 Nash 결과 간의 균형 선택을 위한 시장 설계 레버로 작용하는 방식을 보여준다.
제안 방법
- 각 판매자가 자신의 가격과 실현된 수요만 관찰하는 n명의 판매자가 참여하는 반복 가격 결정 게임과 밴딧 피드백을 모델링한다.
- 두 점 가격 실험(two-point price experimentation)(기준선 대비 소량의 섭동 포함)과 자신의 데이터를 기반으로 하는 선형 회귀 수요 추정기를 도입한다.
- Switchback Linear Demand Learning (SLDL)를 제안한다: 배치 단위 무작위 섭동으로 데이터 수집, OLS 수요 추정, 추정된 수익 최대화 타깃을 향한 부분 가격 업데이트.
- 추측 행렬 A를 통해 라이벌의 가정된 교차 가격 반응을 포착하여 Conjectural Variations (CV) 균형을 정의하고, CV 균형의 1차 조건을 도출한다.
- 상관된 실험이 수요 추정에 편향을 유도하여 경쟁자의 가격 동조를 모방하고, 학습 결과로 CV 균형으로 귀결됨을 보인다.
- 안정성 조건하에서 유한 샘플 보장을 확립한다: 평균 제곱 가격 오차가 로그 인자까지의 차이를 제외하고 감소하는 속도는 T^{-1/2}의 형태를 보인다.
실험 결과
연구 질문
- RQ1다중 판매자 환경에서 간단한 밴딧 피드백 가격 결정 알고리즘이 CV 균형으로 수렴할 수 있는가?
- RQ2가격 실험의 상관 구조가 균형 선택과 가격 책정 결과에 어떤 영향을 미치는가?
- RQ3학습 역학이 수렴하기 위한 수요의 조건은 무엇이며 수렴 속도는 어떠한가?
- RQ4상관된 실험으로 인한 학습 편향이 내부적으로 CV 추측을 어떻게 생성하는가?
- RQ5독립적(비상관) 실험은 Nash 균형으로의 수렴에 어떤 영향을 미치는가?
주요 결과
- 판매자들이 제안된 두 점 실험과 선형 수요 학습을 특정 수요 조건 하에서 따르면 학습 역학은 CV 균형으로 수렴한다.
- 한정된 추측 행렬은 상관된 실험의 통계적 구조에 의해 내생적으로 결정되며 외부 원시적 조건으로 가정되지 않는다.
- 실험이 판매자 간에 비상관적일 때 편향이 사라지고 역학은 표준 Nash 균형으로 수렴한다.
- 본 논문은 표준 모델에서 수렴에 대한 수요의 간단한 충분조건(도함수를 통한)을 제시하고 선형 및 다항 로짓 모델에 대한 명시적 안정성 경계를 제시한다.
- 명시된 안정성 기준 하에 유한 샘플 수렴 보장이 있다: 평균 제곱 가격 오차가 로그 인자를 제외하고 T^{-1/2} 차수로 감소한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.