[论文解读] Canonical correlation coefficients of high-dimensional normal vectors: finite rank case
本文研究了在两个随机向量之间的交叉协方差矩阵具有有限秩的高维设定下,典型相关系数的性质。在维度与样本大小之比几乎必然收敛于 (0,1) 内的常数的条件下,本文证明了最大的样本典型相关系数特征值几乎必然收敛到一个与阈值相关的极限,从而实现了在高维下对真实有限秩相关结构的一致估计。
Consider a normal vector $\mathbf{z}=(\mathbf{x}',\mathbf{y}')'$, consisting of two sub-vectors $\mathbf{x}$ and $\mathbf{y}$ with dimensions $p$ and $q$ respectively. With $n$ independent observations of $\mathbf{z}$ at hand, we study the correlation between $\mathbf{x}$ and $\mathbf{y}$, from the perspective of the Canonical Correlation Analysis, under the high-dimensional setting: both $p$ and $q$ are proportional to the sample size $n$. In this paper, we focus on the case that $Σ_{\mathbf{x}\mathbf{y}}$ is of finite rank $k$, i.e. there are $k$ nonzero canonical correlation coefficients, whose squares are denoted by $r_1\geq\cdots\geq r_k>0$. Under the additional assumptions $(p+q)/n o y\in (0,1)$ and $p/q ot o 1$, we study the sample counterparts of $r_i,i=1,\ldots,k$, i.e. the largest k eigenvalues of the sample canonical correlation matrix $S_{\mathbf{x}\mathbf{x}}^{-1}S_{\mathbf{x}\mathbf{y}}S_{\mathbf{y}\mathbf{y}}^{-1}S_{\mathbf{y}\mathbf{x}}$, namely $λ_1\geq\cdots\geq λ_k$. We show that there exists a threshold $r_c\in(0,1)$, such that for each $i\in\{1,\ldots,k\}$, when $r_i\leq r_c$, $λ_i$ converges almost surely to the right edge of the limiting spectral distribution of the sample canonical correlation matrix, denoted by $d_r$. When $r_i>r_c$, $λ_i$ possesses an almost sure limit in $(d_r,1]$, from which we can recover $r_i$ in turn, thus provide an estimate of the latter in the high-dimensional scenario.
研究动机与目标
- 分析当总体交叉协方差矩阵具有有限秩时,样本典型相关系数的渐近行为。
- 解决在维度 p 和 q 随样本大小 n 比例增长的高维设定下,估计典型相关系数的挑战。
- 在高维渐近框架下,建立最大样本典型相关系数特征值的几乎必然收敛性质。
- 基于其样本对应量,推导一种基于阈值的真正典型相关系数估计程序。
- 为在交叉协方差矩阵中具有有限秩结构的高维多元分析推断提供理论基础。
提出的方法
- 将联合向量 (x, y) 建模为维度分别为 p 和 q 的高维正态随机向量,其中 p, q 与样本大小 n 成比例趋于无穷。
- 将典型相关矩阵定义为 Σ_xx⁻¹Σ_xyΣ_yy⁻¹Σ_yx,其特征值 r_i 表示平方的总体典型相关系数。
- 假设交叉协方差矩阵 Σ_xy 具有有限秩 k,意味着仅有 k 个非零的典型相关系数。
- 分析样本典型相关矩阵 S_xx⁻¹S_xyS_yy⁻¹S_yx,其最大的 k 个特征值 λ_i 是 r_i 的样本对应量。
- 应用随机矩阵理论和有限秩扰动技术,推导样本矩阵的极限谱分布。
- 识别出一个临界阈值 r_c ∈ (0,1),使得当 r_i ≤ r_c 时,λ_i 几乎必然收敛到极限谱分布的右边缘 d_r;当 r_i > r_c 时,λ_i 几乎必然收敛到 (d_r, 1] 内的某个值,该值依赖于 r_i,从而实现对真实 r_i 的一致恢复。
实验结果
研究问题
- RQ1当总体交叉协方差矩阵具有有限秩时,最大样本典型相关系数特征值的渐近行为是什么?
- RQ2在高维设定下,样本典型相关矩阵的极限分布如何依赖于真实的典型相关系数?
- RQ3当 p, q → ∞ 且 p+q ∼ yn 对于 y ∈ (0,1) 时,能否从其样本对应量一致估计真实典型相关系数?
- RQ4阈值 r_c 在决定样本特征值收敛行为中起什么作用?
- RQ5在高维 MANOVA 类型模型中,满足何种条件时,最大样本特征值能够恢复真实典型相关系数?
主要发现
- 当 r_i ≤ r_c 时,样本典型相关系数特征值 λ_i 几乎必然收敛到样本典型相关矩阵极限谱分布的右边缘 d_r。
- 当 r_i > r_c 时,λ_i 几乎必然收敛到区间 (d_r, 1] 内的极限,该极限依赖于 r_i,从而允许对真实 r_i 实现一致恢复。
- 阈值 r_c 严格位于 0 和 1 之间,且依赖于渐近比 y = (p+q)/n ∈ (0,1)。
- λ_i 的极限行为在 r_c 处表现出相变现象,将噪声型特征值与信号型特征值区分开来。
- 研究结果为在高维数据中估计有限秩典型相关结构提供了理论基础,即使当 p 和 q 相对于 n 较大时亦成立。
- 在假设 (p+q)/n → y ∈ (0,1) 且 p/q 不收敛于 1 的条件下,收敛性为几乎必然,确保了在高维框架下的稳定性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。