[论文解读] Likelihood Ratio Test in Multivariate Linear Regression: from Low to High Dimension
本文針對低至高維度設定下的多元線性回歸,發展了一種校正後的概似比檢定(LRT),確立了傳統卡方近似失效的漸近邊界。提出針對 $ p > n $ 的兩步驟程序,驗證其理論性質,並透過模擬與真實資料分析(CNV-GEP關聯)展示改進的檢定效能與準確性。
Multivariate linear regressions are widely used statistical tools in many applications to model the associations between multiple related responses and a set of predictors. To infer such associations, it is often of interest to test the structure of the regression coefficients matrix, and the likelihood ratio test (LRT) is one of the most popular approaches in practice. Despite its popularity, it is known that the classical $χ^2$ approximations for LRTs often fail in high-dimensional settings, where the dimensions of responses and predictors $(m,p)$ are allowed to grow with the sample size $n$. Though various corrected LRTs and other test statistics have been proposed in the literature, the fundamental question of when the classic LRT starts to fail is less studied, an answer to which would provide insights for practitioners, especially when analyzing data with $m/n$ and $p/n$ small but not negligible. Moreover, the power performance of the LRT in high-dimensional data analysis remains underexplored. To address these issues, the first part of this work gives the asymptotic boundary where the classical LRT fails and develops the corrected limiting distribution of the LRT for a general asymptotic regime. The second part of this work further studies the test power of the LRT in the high-dimensional setting. The result not only advances the current understanding of asymptotic behavior of the LRT under alternative hypothesis, but also motivates the development of a power-enhanced LRT. The third part of this work considers the setting with $p>n$, where the LRT is not well-defined. We propose a two-step testing procedure by first performing dimension reduction and then applying the proposed LRT. Theoretical properties are developed to ensure the validity of the proposed method. Numerical studies are also presented to demonstrate its good performance.
研究动机与目标
- 識別傳統卡方近似在多元線性回歸的概似比檢定(LRT)中開始失效的漸近區域。
- 發展LRT統計量的校正極限分配,使其在廣義漸近區域(包含 $ m/n $ 與 $ p/n $ 為小但非可忽略之值的高維設定)中仍具有效性。
- 探討在高維度替代假設下LRT的檢定效能表現,並提出效能增強的變體。
- 針對 $ p > n $ 的情況(標準LRT未定義)提出兩步驟檢定程序,先進行維度降低,再應用校正後的LRT。
- 透過大量模擬與真實基因表現與拷貝數變異資料分析,驗證所提方法。
提出的方法
- 推導在虛無假設下,LRT之傳統 $ \chi^2 $ 近似失效的漸近邊界,識別 $ m, p, n $ 的關鍵比例。
- 提出在廣義漸近區域下LRT統計量的校正極限分配,確保當 $ m, p $ 隨 $ n $ 增長時仍具有效性。
- 針對 $ p > n $ 情境應用兩步驟程序:首先對預測變數進行維度降低(例如透過主成分分析),再於降維後資料上應用校正LRT。
- 利用典型相關與套索迴歸篩選,於檢定前提升對活躍預測變數的選取,以增強檢定效能與覆蓋率。
- 使用蒙地卡羅模擬,針對具有相關性之預測變數($ \rho = 0.7, 0.9 $)與不同信號強度,比較檢定效能與正確覆蓋比例。
- 在真實資料上進行驗證,使用三條染色體資料,$ n_S = 26 $, $ n_T = 63 $, $ J = 2000 $ 等分,並自適應選擇主成分數 $ m_0, p_0 $。
实验结果
研究问题
- RQ1在 $ m, p, n $ 的何種漸近比例下,多元線性回歸中LRT的傳統 $ \chi^2 $ 近似開始失效?
- RQ2如何校正LRT的極限分配,使其在 $ m, p $ 隨 $ n $ 增長的高維設定中仍具有效性?
- RQ3在高維度替代假設下,LRT的檢定效能表現為何?如何加以增強?
- RQ4當 $ p > n $ 時(標準檢定未定義),如何調整LRT?
- RQ5預測變數之間的相關結構如何影響多元回歸中篩選與檢定程序的表現?
主要发现
- 即使樣本數中等,當 $ m/n $ 與 $ p/n $ 為小但非可忽略之值時,LRT的傳統 $ \chi^2 $ 近似即開始失效。
- 所提校正LRT在廣泛的高維度漸近區域中均維持準確的尺寸與有效性。
- 在 $ n=100, p=120, m=5 $ 的模擬中,典型相關基於的篩選在高預測變數相關性下,表現優於套索,無論在檢定效能或正確覆蓋比例上皆更佳。
- 對真正活躍預測變數的選取不足,會持續導致檢定效能降低,凸顯有效篩選的重要性。
- 在CNVs與GEPs的真實資料分析中,相同染色體回歸(如8→8, 17→17)的 $ p $-值顯著低於0.05,支持拒絕虛無假設。
- 對於跨染色體回歸(如22→8),大多數 $ p $-值超過0.05,與無法拒絕虛無假設一致,確認方法在實際應用中的有效性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。