Skip to main content
QUICK REVIEW

[論文レビュー] Nonregular and Minimax Estimation of Individualized Thresholds in High Dimension with Binary Responses

Huijie Feng, Yang Ning|arXiv (Cornell University)|May 26, 2019
Statistical Methods and Inference参考文献 37被引用数 4
ひとこと要約

本稿では、非正則性および計算的に不連続な問題に対処するため、二値応答モデルにおける個別化線形しきい値の高次元推定のための正則化された滑らか化損失手法を提案する。非標準的な誤差率 $$(s\log d/n)^{\beta/(2\beta+1)}$$ を確立し、対数要因を除いて最小最大最適性を示した。また、適応的Lepskiの手法とパスフォローリングアルゴリズムにより、幾何的収束を保証する。

ABSTRACT

Given a large number of covariates $Z$, we consider the estimation of a high-dimensional parameter $θ$ in an individualized linear threshold $θ^T Z$ for a continuous variable $X$, which minimizes the disagreement between $ ext{sign}(X-θ^TZ)$ and a binary response $Y$. While the problem can be formulated into the M-estimation framework, minimizing the corresponding empirical risk function is computationally intractable due to discontinuity of the sign function. Moreover, estimating $θ$ even in the fixed-dimensional setting is known as a nonregular problem leading to nonstandard asymptotic theory. To tackle the computational and theoretical challenges in the estimation of the high-dimensional parameter $θ$, we propose an empirical risk minimization approach based on a regularized smoothed loss function. The statistical and computational trade-off of the algorithm is investigated. Statistically, we show that the finite sample error bound for estimating $θ$ in $\ell_2$ norm is $(s\log d/n)^{β/(2β+1)}$, where $d$ is the dimension of $θ$, $s$ is the sparsity level, $n$ is the sample size and $β$ is the smoothness of the conditional density of $X$ given the response $Y$ and the covariates $Z$. The convergence rate is nonstandard and slower than that in the classical Lasso problems. Furthermore, we prove that the resulting estimator is minimax rate optimal up to a logarithmic factor. The Lepski's method is developed to achieve the adaption to the unknown sparsity $s$ and smoothness $β$. Computationally, an efficient path-following algorithm is proposed to compute the solution path. We show that this algorithm achieves geometric rate of convergence for computing the whole path. Finally, we evaluate the finite sample performance of the proposed estimator in simulation studies and a real data analysis.

研究の動機と目的

  • 二値応答の高次元しきい値推定における計算的不連続性と非正則な漸近的挙動に対処すること。
  • 線形しきい値モデル $$\bm{\theta}^T\bm{Z}$$ における高次元パラメータ $$\bm{\theta}$$ の統計的に最適かつ計算的に効率的な推定手法の開発。
  • 未知のスパarsity $$s$$ と滑らかさ $$\beta$$ の下で、推定量の有限標本誤差境界と最小最大最適性の確立。
  • Lepskiの手法による未知の $$s$$ と $$\beta$$ への適応と、計算における幾何的収束の保証。

提案手法

  • 不連続な符号関数を用いた経験的リスク最小化としてしきい値推定を定式化する。
  • 微分不能な符号関数の代わりに、最適化が容易になる正則化された滑らか化損失関数を導入する。
  • 滑らかさパラメータ $$\delta \to 0$$ の下でFisherの整合性を確立し、真のリスク最小化器に収束することを保証する。
  • 条件付き密度 $$X$$ が $$Y$$ と $$\bm{Z}$$ を与えた下での滑らかさ $$\beta$$ の下で、$$\ell_2$$ 誤差境界 $$(s\log d/n)^{\beta/(2\beta+1)}$$ を導出する。
  • $$s$$ と $$\beta$$ の事前知識が不要な状況下で、Lepskiの手法を用いてチューニングパラメータを適応的に選択する。
  • 全解パスを効率的に計算するため、幾何的収束率を達成するパスフォローリングアルゴリズムを開発する。

実験結果

リサーチクエスチョン

  • RQ1非正則性下で、計算的に実行可能かつ統計的に最適な高次元しきい値推定手法を、二値応答に対して開発可能か?
  • RQ2未知の滑らかさ $$\beta$$ とスパarsity $$s$$ の下で、$$\ell_2$$ ノルムにおける $$\bm{\theta}$$ の推定の最適収束速度は何か?
  • RQ3高次元しきい値推定において、未知の $$s$$ と $$\beta$$ への適応はどのように達成できるか?
  • RQ4非凸的かつ非滑らかな経験的リスク最小化問題は、正則化された滑らか化損失アプローチによって効果的に解けるか?
  • RQ5提案手法は、高次元的かつ非正則な設定下で、対数要因を除いて最小最大最適性を達成するか?

主な発見

  • 提案された推定量は、$$(s\log d/n)^{\beta/(2\beta+1)}$$ の $$\ell_2$$ 誤差境界を達成し、これは非標準的で、古典的なLassoレートより遅い。
  • 収束速度は、対数要因を除いて最小最大最適であり、手法の理論的最適性を確立する。
  • Lepskiの手法により、$$s$$ と $$\beta$$ の事前知識がなくても、未知のスパarsityと滑らかさへの適応が可能である。
  • パスフォローリングアルゴリズムは幾何的収束率を達成し、全解パスの効率的計算を保証する。
  • ChAMP試験からのシミュレーションと実データ解析により、有限標本における性能とさまざまな変数選択パターンへのロバストネスが確認された。
  • SVM やロジスティック回帰と比較して、臨床的に関連のある変数(例:KSymp_3mo および SF36Soc_6mo)が一貫して負の係数で選択された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。