Skip to main content
QUICK REVIEW

[論文レビュー] Balancing Covariates in Randomized Experiments with the Gram-Schmidt Walk Design

Venkat, Prayaag|arXiv (Cornell University)|Nov 8, 2019
Advanced Causal Inference Techniques参考文献 42被引用数 13
ひとこと要約

本稿では、調整可能なパラメータを用いて共変量のバランスとロバストネスを制御する、新しい実験設計であるグラム・シュミット・ウォーク(GSW)設計を導入する。この設計は漸近的にほぼ完全な共変量のバランスを達成しつつ、有限標本におけるロバストネスを維持するため、保守的な信頼区間を伴う漸近的に有効な推論を可能にする。

ABSTRACT

In this paper, we initiate the study of the algorithmic problem of certifying lower bounds on the discrepancy of random matrices: given an input matrix A ∈ ℝ^{m × n}, output a value that is a lower bound on disc(A) = min_{x ∈ {± 1}ⁿ} ‖Ax‖_∞ for every A, but is close to the typical value of disc(A) with high probability over the choice of a random A. This problem is important because of its connections to conjecturally-hard average-case problems such as negatively-spiked PCA [Afonso S. Bandeira et al., 2020], the number-balancing problem [Gamarnik and Kızıldağ, 2021] and refuting random constraint satisfaction problems [Prasad Raghavendra et al., 2017]. We give the first polynomial-time algorithms with non-trivial guarantees for two main settings. First, when the entries of A are i.i.d. standard Gaussians, it is known that disc(A) = Θ (√n2^{-n/m}) with high probability [Karthekeyan Chandrasekaran and Santosh S. Vempala, 2014; Aubin et al., 2019; Paxton Turner et al., 2020] and that super-constant levels of the Sum-of-Squares SDP hierarchy fail to certify anything better than disc(A) ≥ 0 when m < n - o(n) [Mrinalkanti Ghosh et al., 2020]. In contrast, our algorithm certifies that disc(A) ≥ exp(-O(n²/m)) with high probability. As an application, this formally refutes a conjecture of Bandeira, Kunisky, and Wein [Afonso S. Bandeira et al., 2020] on the computational hardness of the detection problem in the negatively-spiked Wishart model. Second, we consider the integer partitioning problem: given n uniformly random b-bit integers a₁, …, a_n, certify the non-existence of a perfect partition, i.e. certify that disc(A) ≥ 1 for A = (a₁, …, a_n). Under the scaling b = α n, it is known that the probability of the existence of a perfect partition undergoes a phase transition from 1 to 0 at α = 1 [Christian Borgs et al., 2001]; our algorithm certifies the non-existence of perfect partitions for some α = O(n). We also give efficient non-deterministic algorithms with significantly improved guarantees, raising the possibility that the landscape of these certification problems closely resembles that of e.g. the problem of refuting random 3SAT formulas in the unsatisfiable regime. Our algorithms involve a reduction to the Shortest Vector Problem and employ the Lenstra-Lenstra-Lovász algorithm.

研究の動機と目的

  • 実験設計における共変量のバランスとロバストネスの根本的トレードオフに対処する。
  • 最悪ケースの平均二乗誤差を用いてロバストネスを形式化し、共変量の線形関数を用いてバランスを形式化する。
  • 実験者が1つのパラメータでバランスとロバストネスのトレードオフを調整できる設計を開発する。
  • 漸近的正規性を保証し、有効な推論を可能にする保守的な分散推定器を提供する。
  • 大標本において、この設計がバランスとロバストネスのトレードオフを回避し、完全なバランスを達成しつつロバストネスの損失を小さくできることを示す。

提案手法

  • 実験設計を分布的乖離問題として定式化し、実験設計をアルゴリズム的乖離に翻訳する。
  • バンサルら(2019)のグラム–シュミット・ウォークアルゴリズムを改変し、確率的処置割り当て分布を生成する。
  • ロバストネスパラメータϕを用いて、ホルヴィッツ–トムプソン推定量の最悪ケースの平均二乗誤差を制限する。
  • すべての共変量の線形関数が同時にバランスされるように設計し、リッジ回帰損失から導かれる境界を適用する。
  • 乖離理論を用いて、有限標本における平均二乗誤差と標本分布の尾部に関する境界を導出する。
  • 漸近的に有効な信頼区間を保証する保守的で一貫性のある分散推定器を提供する。

実験結果

リサーチクエスチョン

  • RQ1実験設計におけるバランスとロバストネスのトレードオフを形式的に定量化し、その乗り越え方をどのように設計できるか。
  • RQ21つの設計が、多様な共変量構成において高い共変量バランスと十分なロバストネスを同時に達成できるか。
  • RQ3GSW設計下でのホルヴィッツ–トムプソン推定量の漸近的挙動はいかなるものか。
  • RQ4GSW設計は保守的な信頼区間を伴う漸近的に有効な推論を可能にするか。
  • RQ5バランス、ロバストネス、平均二乗誤差の観点から、GSW設計は再ランダマイズやマッチドペア設計と比べてどのように異なるか。

主な発見

  • グラム・シュミット・ウォーク設計では、潜在的アウトカムを共変量に対して暗黙のリッジ回帰として扱った際の損失によって、ホルヴィッツ–トムプソン推定量の平均二乗誤差が制限される。
  • 大標本では、GSW設計はすべての共変量の線形関数について漸近的に完全なバランスを達成し、バランスとロバストネスのトレードオフを回避する。
  • n = 2960の場合、ϕ = 0.01のGSW設計では、バランス演算子ノルムが再ランダマイズの約57倍小さく、マッチドペア設計の約41倍小さかった。
  • ϕ = 0.01のGSW設計下では、アウトカムDについて、平均二乗誤差がマッチドペア設計の31倍以上小さかった。
  • GSW設計(ϕ = 0.01)における主な信頼区間の幅は、ベルヌーイ分布および完全ランダマイズのそれよりも約10–14%狭く、n ≥ 296の条件下で名目水準を超えるカバレッジを示した。
  • 保守的な分散推定器により、漸近的に有効な信頼区間が保証され、有限標本の境界に基づく代替区間は高いカバレッジを示したが、幅が広がった。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。