Skip to main content
QUICK REVIEW

[論文レビュー] Two-sample testing in non-sparse high-dimensional linear models

Yinchu Zhu, Jelena Bradić|arXiv (Cornell University)|Oct 14, 2016
Statistical Methods and Inference参考文献 5被引用数 6
ひとこと要約

本稿では、回帰係数のスパarsityを仮定しない高次元線形モデルにおける二標本検定フレームワークであるTIERSを提案する。二つの標本を畳み込みることで等しい回帰係数の帰無仮説をモーメント条件に変換し、自己正規化を用いることで、弱いモーメントおよび尾部条件のもとでも、$ p \to \infty $ かつ $ n \to \infty $ かつ $ p/n \to \infty $ の場合でも、タイプIエラーを安定に制御する。また、自動的・適応的Dantzig選択子(ADDS)と高次元プラグイン臨界値近似法を導入する。

ABSTRACT

In analyzing high-dimensional models, sparsity of the model parameter is a common but often undesirable assumption. In this paper, we study the following two-sample testing problem: given two samples generated by two high-dimensional linear models, we aim to test whether the regression coefficients of the two linear models are identical. We propose a framework named TIERS (short for TestIng Equality of Regression Slopes), which solves the two-sample testing problem without making any assumptions on the sparsity of the regression parameters. TIERS builds a new model by convolving the two samples in such a way that the original hypothesis translates into a new moment condition. A self-normalization construction is then developed to form a moment test. We provide rigorous theory for the developed framework. Under very weak conditions of the feature covariance, we show that the accuracy of the proposed test in controlling Type I errors is robust both to the lack of sparsity in the features and to the heavy tails in the error distribution, even when the sample size is much smaller than the feature dimension. Moreover, we discuss minimax optimality and efficiency properties of the proposed test. Simulation analysis demonstrates excellent finite-sample performance of our test. In deriving the test, we also develop tools that are of independent interest. The test is built upon a novel estimator, called Auto-aDaptive Dantzig Selector (ADDS), which not only automatically chooses an appropriate scale of the error term but also incorporates prior information. To effectively approximate the critical value of the test statistic, we develop a novel high-dimensional plug-in approach that complements the recent advances in Gaussian approximation theory.

研究の動機と目的

  • 回帰係数が非スパースである場合に、高次元二標本検定のための体系的でない推論手法の欠如に対処すること。
  • 特徴量の共分散構造や誤差分布に対する弱い仮定のもとでも、特に重尾分布の下でも、正確なタイプIエラー制御を維持する検定を開発すること。
  • $ p \to \infty $, $ n \to \infty $, および $ p/n \to \infty $ の下でも、真の回帰パラメータのスパarsityを要件としない、耐性のある手法を構築すること。
  • 誤差分散のスケールを適応的に選択し、事前情報を取り入れる新しい推定量である自動的・適応的Dantzig選択子(ADDS)を導入すること。
  • 最近のガウス近似理論を補完する高次元プラグイン法を用いて、検定統計量の臨界値を近似する手法を開発すること。

提案手法

  • TIERSは、二つの標本の畳み込みを構築することで、元の二標本仮説 $ H_0: \beta_A = \beta_B $ を新たなモーメント条件に変換する。
  • 変換されたデータに対して自己正規化手順を適用し、帰無仮説の下で漸近的に分布不変となるピボット統計量を構成する。
  • 自動的・適応的Dantzig選択子(ADDS)は、誤差分散のスケールを適応的に選択し、事前知識を取り入れる新しい推定量として導入される。
  • 検定統計量の臨界値は、高次元プラグイン法を用いて近似され、ブートストラップやリサンプリングを回避する。
  • 理論的分析では、高次元確率的ベクトルの最大値の逸脱を制御するため、反濃度不等式とサブガウス型尾部バウンドに依拠する。
  • 従属確率的ベクトルの最大値の収束および条件付きガウス分布下での反濃度に関する補題を用い、検定の耐性を確立する。

実験結果

リサーチクエスチョン

  • RQ1回帰係数のスパarsity仮定に依存しない、高次元線形モデルにおける二標本検定手順を開発することは可能か?
  • RQ2特徴量数 $ p $ が標本サイズ $ n $ よりも著しく大きい場合、特に非スパースモデル下でも、タイプIエラーの正確な制御は可能か?
  • RQ3重尾誤差および非スパース設計行列が、高次元二標本検定の妥当性に与える影響は何か?
  • RQ4リサンプリングを用いずに、高次元設定下でプラグイン法により検定統計量の臨界値を近似できるか?
  • RQ5提案された検定は、非スパースな高次元モデル下でミニマックス最適性または効率性を達成するか?

主な発見

  • TIERSは、真の回帰係数が密であっても、特徴量の共分散構造に対する非常に弱い条件下でも、タイプIエラーを正確に制御する。
  • 重尾誤差分布に対しても耐性を示し、サブガウス型または軽尾仮定を必要としない。
  • 自動的・適応的Dantzig選択子(ADDS)は、誤差分散のスケールを自動的に選択し、事前情報を取り入れることで推定の安定性を向上させる。
  • 臨界値の高次元プラグイン近似法は、良好な有限標本性能を達成し、計算コストの高いリサンプリングを回避する。
  • 理論的分析により、回帰パラメータの非スパarsityおよび非スパース設計行列に対しても、検定が耐性を示すことが示された。$ p/n \to \infty $ の場合でも同様である。
  • シミュレーション研究を通じて、適切な正則性条件下で、この手法がミニマックス最適性および効率性を達成することが確認された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。