[Paper Review] Repeated Bilateral Trade Against a Smoothed Adversary
This paper studies repeated bilateral trade under a σ-smooth adversary, where seller and buyer valuations are adaptively chosen. It proposes a novel learning algorithm that achieves T^{3/4} regret under partial feedback when different prices can be set for buyers and sellers, proving this rate is optimal via a surprising lower bound, which surpasses the standard √T and T^{2/3} regimes in partial feedback settings.
We study repeated bilateral trade where an adaptive $σ$-smooth adversary generates the valuations of sellers and buyers. We provide a complete characterization of the regret regimes for fixed-price mechanisms under different feedback models in the two cases where the learner can post either the same or different prices to buyers and sellers. We begin by showing that the minimax regret after $T$ rounds is of order $\sqrt{T}$ in the full-feedback scenario. Under partial feedback, any algorithm that has to post the same price to buyers and sellers suffers worst-case linear regret. However, when the learner can post two different prices at each round, we design an algorithm enjoying regret of order $T^{3/4}$ ignoring log factors. We prove that this rate is optimal by presenting a surprising $T^{3/4}$ lower bound, which is the main technical contribution of the paper.
Motivation & Objective
- To study online learning in repeated bilateral trade under a σ-smooth adversary, which generalizes i.i.d. and worst-case settings.
- To characterize the minimax regret for fixed-price mechanisms under different feedback models.
- To determine whether sublinear regret is achievable when the learner can set distinct prices for buyers and sellers under partial feedback.
- To establish the optimality of the T^{3/4} regret rate in this setting.
Proposed method
- The paper introduces a smoothed adversary model where the joint distribution of seller and buyer valuations is σ-smooth, ensuring limited adversarial variation over time.
- It designs a variant of the Exp3 algorithm, called Blind-Exp3, which uses a grid-based exploration strategy and importance-weighted rewards to estimate gains from trade.
- The algorithm maintains weights over a discretized price grid and updates them using exponential weights based on estimated gains, with exploration controlled by a parameter γ.
- A key technical component is the use of a bounded reward estimator ˆrt(i) that combines exploitation and exploration to ensure concentration of the estimated gain.
- The analysis leverages logarithmic and exponential inequalities under parameter constraints (2ηK/γ ≤ 1) to derive regret bounds.
- A novel T^{3/4} lower bound is proven using information-theoretic arguments, establishing tightness of the upper bound.
Experimental results
Research questions
- RQ1Can sublinear regret be achieved in repeated bilateral trade when the learner faces a smoothed adversary and only partial feedback is available?
- RQ2What is the optimal regret rate when the learner can set different prices for buyers and sellers under partial feedback?
- RQ3Is the T^{3/4} regret rate tight, or can it be improved further under the smoothed adversary model?
- RQ4How does the minimax regret in this setting compare to classical partial feedback models like partial monitoring or feedback graphs?
Key findings
- The minimax regret for fixed-price mechanisms is Θ(√T) under full feedback, matching standard online learning results.
- When the same price must be posted to both buyers and sellers, any algorithm suffers worst-case linear regret under partial feedback.
- When different prices can be set for buyers and sellers, the proposed Blind-Exp3 algorithm achieves regret of order T^{3/4} ignoring logarithmic factors.
- The paper establishes a T^{3/4} lower bound, proving that this rate is optimal and cannot be improved.
- The result reveals a new regret regime that surpasses the √T vs. T^{2/3} dichotomy observed in other partial feedback models.
- The analysis shows that the smoothed adversary model enables sublinear regret even with partial feedback, extending learnability beyond i.i.d. assumptions.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.