Skip to main content
QUICK REVIEW

[Paper Review] Distributionally balanced sampling designs

Anton Grafström, Wilmer Prentius|arXiv (Cornell University)|Mar 12, 2026
Optimal Experimental Design Methods0 citations
TL;DR

Introduces Distributionally Balanced Designs (DBD), a probability sampling class that minimizes the energy distance between the sample and population auxiliary distributions via a circular, optimized ordering and random contiguous block selection to improve distributional representativeness and estimator variance.

ABSTRACT

We propose Distributionally Balanced Designs (DBD), a new class of probability sampling designs that target representativeness at the level of the full auxiliary distribution rather than selected moments. In disciplines such as ecology, forestry, and environmental sciences, where field data collection is expensive, maximizing the information extracted from a limited sample is critical. More precisely, DBD can be viewed as minimum discrepancy designs that minimize the expected discrepancy between the sample and population auxiliary distributions. The key idea is to construct samples whose empirical auxiliary distribution closely matches that of the population. We present a first implementation of DBD based on an optimized circular ordering of the population, combined with random selection of a contiguous block of units. The ordering is chosen to minimize the design-expected energy distance, a discrepancy measure that captures differences between distributions beyond low-order moments. This criterion promotes strong spatial spread, and yields low variance for Horvitz-Thompson estimators of totals of functions that vary smoothly with respect to auxiliaries. Simulation results show that approximate DBD achieves better distributional fit than state-of-the-art methods such as the local pivotal and local cube designs. Hence, DBD can improve the reliability of estimates from costly field data, making distributional balancing effective for constructing representative surveys in resource-constrained applications.

Motivation & Objective

  • Motivate the need for distribution-wide representativeness beyond means or spatial spread.
  • Propose a formal framework (DBD) to minimize distributional discrepancy between sample and population.
  • Develop an optimization-based construction (circular ordering + contiguous block) to approximate distributional balance.
  • Provide variance estimation guidance and assess performance via simulations and real data.
  • Offer scalable implementation guidance and discuss applicability beyond traditional survey sampling.

Proposed method

  • Define Distributionally Balanced Designs (DBD) as designs minimizing the expected energy distance between sample and population auxiliary distributions.
  • Adopt energy distance (a form of Maximum Mean Discrepancy) as the discrepancy measure to capture all moments.
  • Restrict design class to equal-probability designs formed by circular permutations and random starting points.
  • Use simulated annealing to optimize the circular order of the population to minimize the average sample-population energy distance.
  • Leverage a fast O(n) update for objective evaluation per swap to enable efficient optimization.
  • Provide a local-mean variance estimator for variance estimation suited to highly spread samples.

Experimental results

Research questions

  • RQ1How can sampling designs be constructed so that the sample's auxiliary distribution closely matches the population's distribution?
  • RQ2Does optimizing for distributional fit (energy distance) yield improved variance properties for Horvitz-Thompson estimators under smooth target functions?
  • RQ3How does DBD compare with state-of-the-art methods (LPM, LCUBE, SRS) in terms of distributional fit, spatial spread, and local balance across varying dimensionality of auxiliaries?
  • RQ4Is the circular DBD scalable to larger populations, and can a block/stratified version preserve variance reductions?

Key findings

  • DBD achieves better distributional fit (lower mean energy distance) than local pivotal and local cube designs across dimensions.
  • Optimized circular sequencing yields strong spatial spread while maintaining equal inclusion probabilities.
  • DBD shows superior balance-related metrics (LB and BD) compared with competing designs, especially in lower dimensions.
  • Variance estimation with a local-mean approach adapts to the smoothness structure of the target function and is stable under DBD.
  • As sample size grows, the distributional advantages of DBD compound, with faster decay in balance deviation than SRS.
  • On real data (Meuse), circular DBD provides the lowest energy distance and more accurate estimates for auxiliary and target variables, with conservative coverage.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.