Skip to main content
QUICK REVIEW

[Paper Review] Load Balancing in the Non-Degenerate Slowdown Regime

Varun Gupta, Neil Walton|arXiv (Cornell University)|Jul 6, 2017
Advanced Queuing Theory Analysis34 references6 citations
TL;DR

This paper introduces a novel analysis of the Join-the-Shortest-Queue (JSQ) load balancing policy in the Non-Degenerate Slowdown (NDS) regime, a many-server asymptotic framework where the number of spare servers remains fixed as system size grows. It derives a diffusion approximation for JSQ, quantifies a 15% performance gap between JSQ and centralized queueing, and proposes a new low-overhead policy, Idle-One-First (I1F), that matches JSQ’s asymptotic performance with reduced communication overhead.

ABSTRACT

We analyse Join-the-Shortest-Queue in a contemporary scaling regime known as the Non-Degenerate Slowdown regime. Join-the-Shortest-Queue (JSQ) is a classical load balancing policy for queueing systems with multiple parallel servers. Parallel server queueing systems are regularly analysed and dimensioned by diffusion approximations achieved in the Halfin-Whitt scaling regime. However, when jobs must be dispatched to a server upon arrival, we advocate the Non-Degenerate Slowdown regime (NDS) to compare different load-balancing rules. In this paper we identify novel diffusion approximation and timescale separation that provides insights into the performance of JSQ. We calculate the price of irrevocably dispatching jobs to servers and prove this to within 15% (in the NDS regime) of the rules that may manoeuvre jobs between servers. We also compare ours results for the JSQ policy with the NDS approximations of many modern load balancing policies such as Idle-Queue-First and Power-of-$d$-choices policies which act as low information proxies for the JSQ policy. Our analysis leads us to construct new rules that have identical performance to JSQ but require less communication overhead than power-of-2-choices.

Motivation & Objective

  • To analyze the performance of Join-the-Shortest-Queue (JSQ) in the Non-Degenerate Slowdown (NDS) regime, a non-standard asymptotic framework that better distinguishes load-balancing policies than traditional regimes.
  • To quantify the performance gap between optimal dispatch (JSQ) and optimal pooling (centralized queueing), particularly in terms of mean response time.
  • To design a low-information dispatch rule that achieves JSQ-like performance with significantly reduced communication overhead compared to existing policies like Power-of-2-choices.
  • To demonstrate that the NDS regime enables meaningful differentiation between load-balancing policies—unlike Heavy Traffic or Halfin-Whitt regimes—where such distinctions vanish.

Proposed method

  • Applies a diffusion approximation to the JSQ policy in the NDS regime, defined by a fixed number of spare servers as the number of servers $k \to \infty$.
  • Uses stochastic averaging principles, time-scale separation, and coupling techniques to analyze the limiting behavior of the system under JSQ and related policies.
  • Derives the stationary distribution of the diffusion limit for JSQ and compares it with those of other policies using stochastic dominance arguments based on drift functions.
  • Applies the Martingale Functional Central Limit Theorem to justify the convergence of the scaled queueing processes to a diffusion process.
  • Introduces and analyzes the Idle-One-First (I1F) policy, which prioritizes idle servers and those with one job, showing it matches JSQ’s stationary distribution asymptotically.
  • Compares the performance of JSQ, I1F, IQF, and Power-of-$d$-choices policies in the NDS regime using stochastic ordering and invariant measure analysis.

Experimental results

Research questions

  • RQ1How does the performance of the Join-the-Shortest-Queue (JSQ) policy compare to centralized queueing in the Non-Degenerate Slowdown (NDS) regime?
  • RQ2Can a low-overhead load-balancing policy achieve asymptotic performance equivalent to JSQ while minimizing communication cost?
  • RQ3Why is the NDS regime better suited than traditional asymptotic frameworks (e.g., Halfin-Whitt or Heavy Traffic) for distinguishing between load-balancing policies?
  • RQ4What is the exact performance gap between JSQ and centralized queueing in terms of mean response time under the NDS regime?
  • RQ5How does the Idle-One-First (I1F) policy compare to JSQ and other low-information policies in terms of stationary distribution and system performance?

Key findings

  • In the NDS regime, the mean response time of JSQ is at most 15% higher than that of the centralized $M/M/k$ queue, quantifying the cost of not being able to pool or jockey jobs.
  • The Idle-One-First (I1F) policy achieves the same asymptotic stationary distribution as JSQ, implying identical mean response time, despite requiring less communication than Power-of-2-choices.
  • The NDS regime enables a clear distinction between load-balancing policies: in contrast, the Halfin-Whitt and Heavy Traffic regimes render JSQ and centralized queueing indistinguishable.
  • The stationary distribution of the JSQ diffusion in the NDS regime is characterized via a Langevin diffusion with a specific drift function, and its invariant measure is derived explicitly.
  • The performance ordering $\pi_{CQ} \leq_{st} \pi_{JSQ} =_{st} \pi_{I1F} \leq_{st} \pi_{IQF}$ holds due to stochastic dominance of drift terms, confirming the relative performance of the policies.
  • The analysis confirms that the Idle-Queue-First (IQF) policy can have up to 100% higher mean response time than centralized queueing, highlighting the need for refined low-overhead alternatives like I1F.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.