Skip to main content
QUICK REVIEW

[Paper Review] Online Ad Procurement in Non-stationary Autobidding Worlds

Jason Cheuk Nam Liang, Haihao Lu|arXiv (Cornell University)|Jul 10, 2023
Advanced Bandit Algorithms ResearchDecision Sciences3 citations
TL;DR

This paper proposes a primal-dual online learning algorithm for dynamic optimization of ad campaign levers in non-stationary autobidding environments with bandit feedback and uncertain long-term constraints. It achieves low regret across diverse non-stationary worlds—stochastic, adversarial, corrupted, periodic, and ergodic—without prior knowledge of the underlying data-generating process, enabling robust advertiser decision-making in real-world online advertising platforms.

ABSTRACT

Today's online advertisers procure digital ad impressions through interacting with autobidding platforms: advertisers convey high level procurement goals via setting levers such as budget, target return-on-investment, max cost per click, etc.. Then ads platforms subsequently procure impressions on advertisers' behalf, and report final procurement conversions (e.g. click) to advertisers. In practice, advertisers may receive minimal information on platforms' procurement details, and procurement outcomes are subject to non-stationary factors like seasonal patterns, occasional system corruptions, and market trends which make it difficult for advertisers to optimize lever decisions effectively. Motivated by this, we present an online learning framework that helps advertisers dynamically optimize ad platform lever decisions while subject to general long-term constraints in a realistic bandit feedback environment with non-stationary procurement outcomes. In particular, we introduce a primal-dual algorithm for online decision making with multi-dimension decision variables, bandit feedback and long-term uncertain constraints. We show that our algorithm achieves low regret in many worlds when procurement outcomes are generated through procedures that are stochastic, adversarial, adversarially corrupted, periodic, and ergodic, respectively, without having to know which procedure is the ground truth. Finally, we emphasize that our proposed algorithm and theoretical results extend beyond the applications of online advertising.

Motivation & Objective

  • To address the challenge of dynamic, high-dimensional lever optimization in online advertising under limited feedback and non-stationary market conditions.
  • To develop a unified algorithm that performs well across multiple types of non-stationary procurement environments without prior knowledge of the data-generating process.
  • To ensure long-term constraint satisfaction (e.g., budget, ROI) while operating under bandit feedback where only outcomes—not full function values—are observed.
  • To model real-world ad procurement as an online constrained optimization problem with function-valued bandit feedback and uncertain constraints.
  • To extend theoretical guarantees to non-stationary worlds including stochastic, adversarial, and periodically varying outcomes.

Proposed method

  • Formulates the ad procurement problem as online constrained optimization with bandit feedback and uncertain constraints, modeling lever decisions as multi-dimensional variables.
  • Introduces a primal-dual algorithm that maintains dual variables for constraints and uses projected gradient updates with adaptive step-sizes.
  • Employs a randomized exploration strategy via exponential weights over candidate lever configurations to balance exploration and exploitation.
  • Uses Hoeffding’s lemma and convex analysis to bound regret in terms of function variation, constraint violation, and decision diameter.
  • Derives regret bounds that scale sublinearly with time, independent of the specific non-stationary process, under general assumptions on function smoothness and bounded gradients.
  • Applies a smoothing technique using uniform sampling over the unit sphere to estimate gradients in the absence of full function feedback.

Experimental results

Research questions

  • RQ1Can a single online learning algorithm achieve low regret across diverse non-stationary procurement environments without prior knowledge of the data-generating process?
  • RQ2How can advertisers dynamically optimize multiple ad campaign levers under long-term budget and ROI constraints with only bandit feedback?
  • RQ3What theoretical guarantees can be provided for online optimization under uncertain, time-varying constraints and partial feedback?
  • RQ4How can primal-dual methods be adapted to handle function-valued feedback and multi-dimensional decision variables in non-stationary settings?
  • RQ5Can the algorithm maintain low regret in adversarial, stochastic, and periodically varying environments simultaneously?

Key findings

  • The proposed algorithm achieves sublinear regret in time T, with regret upper bounded by O(√T) under general non-stationary conditions including stochastic, adversarial, and periodic outcomes.
  • The algorithm ensures low constraint violation over time, with theoretical bounds on cumulative constraint violations that scale sublinearly with T.
  • The regret bound holds without prior knowledge of the non-stationary process type, making the algorithm robust across diverse real-world ad market dynamics.
  • The algorithm's performance is independent of the specific nature of non-stationarity—whether the outcomes are i.i.d., adversarial, or periodically varying.
  • Theoretical analysis confirms that the algorithm maintains low regret even under adversarial corruptions and ergodic non-stationary processes.
  • The method extends beyond online advertising, offering a general framework for online optimization under uncertainty and partial feedback.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.