Skip to main content
QUICK REVIEW

[論文レビュー] Likelihood Ratio Test in Multivariate Linear Regression: from Low to High Dimension

Yinqiu He, Tiefeng Jiang|arXiv (Cornell University)|Dec 17, 2018
Random Matrices and Applications参考文献 35被引用数 5
ひとこと要約

本稿は、低次元から高次元の設定にわたり、多変量線形回帰における補正済み尤度比検定(LRT)を開発し、古典的なカイ二乗近似が失敗する漸近的境界を確立する。$ p > n $ の設定に対して二段階手順を提案し、その理論的性質を検証し、シミュレーションおよびCNV-GEP関連性の実データ解析を通じて、検出力と精度の向上を示す。

ABSTRACT

Multivariate linear regressions are widely used statistical tools in many applications to model the associations between multiple related responses and a set of predictors. To infer such associations, it is often of interest to test the structure of the regression coefficients matrix, and the likelihood ratio test (LRT) is one of the most popular approaches in practice. Despite its popularity, it is known that the classical $χ^2$ approximations for LRTs often fail in high-dimensional settings, where the dimensions of responses and predictors $(m,p)$ are allowed to grow with the sample size $n$. Though various corrected LRTs and other test statistics have been proposed in the literature, the fundamental question of when the classic LRT starts to fail is less studied, an answer to which would provide insights for practitioners, especially when analyzing data with $m/n$ and $p/n$ small but not negligible. Moreover, the power performance of the LRT in high-dimensional data analysis remains underexplored. To address these issues, the first part of this work gives the asymptotic boundary where the classical LRT fails and develops the corrected limiting distribution of the LRT for a general asymptotic regime. The second part of this work further studies the test power of the LRT in the high-dimensional setting. The result not only advances the current understanding of asymptotic behavior of the LRT under alternative hypothesis, but also motivates the development of a power-enhanced LRT. The third part of this work considers the setting with $p>n$, where the LRT is not well-defined. We propose a two-step testing procedure by first performing dimension reduction and then applying the proposed LRT. Theoretical properties are developed to ensure the validity of the proposed method. Numerical studies are also presented to demonstrate its good performance.

研究の動機と目的

  • 多変量線形回帰における尤度比検定(LRT)の古典的カイ二乗近似が失敗し始める漸近的領域を特定すること。
  • 一般の漸近的領域、特に$ m/n $および$ p/n $が小さく非無視的であるが高次元設定を含む、古典的カイ二乗近似が失敗する領域を特定すること。
  • 高次元代替仮説下におけるLRTの検出力性能を調査し、検出力を向上させる変種を提案すること。
  • $ p > n $ の場合、標準LRTが定義されないため、次元削減の後に補正LRTを適用する二段階検定手順を提案すること。
  • 広範なシミュレーションおよび遺伝子発現とコピー数変異の実データ解析を通じて、提案手法の妥当性を検証すること。

提案手法

  • 帰無仮説下におけるLRTの古典的$ \chi^2 $近似の失敗の漸近的境界を導出し、$ m, p, n $ の重要なスケーリングを同定する。
  • 一般の漸近的領域におけるLRT統計量の補正された極限分布を提案し、$ m, p $ が $ n $ とともに増加する場合でも有効性を保証する。
  • $ p > n $ の場合の二段階手順を適用:まず予測子に対して次元削減(例:主成分分析)を実施し、その後、低次元化されたデータに対して補正LRTを適用する。
  • 検出力とカバレッジの向上を図るため、事前スクリーニングとして正準相関およびラassoベースのスクリーニングを用いる。
  • 相関のある予測子($ \rho = 0.7, 0.9 $)と変動する信号サイズを想定したモンテカルロシミュレーションを実施し、検出力と正しい信頼区間カバー率を比較する。
  • 適応的に選択された主成分($ m_0, p_0 $)を用いて、3本の染色体からの実データを検証し、$ n_S = 26 $, $ n_T = 63 $, $ J = 2000 $ スプリットを用いる。

実験結果

リサーチクエスチョン

  • RQ1多変量線形回帰におけるLRTの古典的$ \chi^2 $近似が失敗し始める$ m, p, n $ の漸近的スケーリングは何か?
  • RQ2高次元設定($ m, p $ が $ n $ とともに増加)においても有効性を保つように、LRTの極限分布をどのように補正できるか?
  • RQ3高次元代替仮説下におけるLRTの検出力行動は何か? また、どのようにして検出力を向上させられるか?
  • RQ4標準LRTが定義されない$ p > n $ の場合、LRTはどのように適合できるか?
  • RQ5予測子間の相関構造が、多変量回帰におけるスクリーニングおよび検定手順の性能にどのように影響するか?

主な発見

  • 古典的$ \chi^2 $近似は、$ m/n $ および $ p/n $ が小さく非無視的であっても、中程度のサンプルサイズでも失敗する。
  • 提案された補正LRTは、広範な高次元漸近的領域において、正確なサイズと有効性を維持する。
  • 相関が高い予測子を想定したシミュレーション($ n=100, p=120, m=5 $)において、正準相関ベースのスクリーニングはラassoを上回り、検出力と正しいカバー率の両面で優れた性能を示した。
  • 真の活性予測子の過小選択は一貫して検出力を低下させ、効果的なスクリーニングの重要性を浮き彫りにした。
  • CNVsとGEPsの実データ解析において、同じ染色体間の回帰(例:8→8, 17→17)の$ p $-値は有意に0.05未満であり、帰無仮説の棄却を支持する。
  • 相互染色体間の回帰(例:22→8)では、大多数の$ p $-値が0.05を超えており、帰無仮説を棄却できないことと整合的であり、実用的妥当性を裏付けた。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。