Skip to main content
QUICK REVIEW

[论文解读] High dimensionality: The latest challenge to data analysis

Ana M. Pires, João A. Branco|arXiv (Cornell University)|Feb 12, 2019
Gene expression and cancer classification参考文献 15被引用 13
一句话总结

本文揭示了当变量数 $ p $ 接近或超过观测数 $ n $ 时,在高维数据分析中存在基本的几何与数学悖论,证明了传统多元方法——尤其是依赖马氏距离与协方差结构的方法——在数学上变得无效。研究证明,在 $ p \geq n-1 $ 的情况下,所有数据点在投影中变为等距,导致基于距离的推断失去意义。

ABSTRACT

The advent of modern technology, permitting the measurement of thousands of characteristics simultaneously, has given rise to floods of data characterized by many large or even huge datasets. This new paradigm presents extraordinary challenges to data analysis and the question arises: how can conventional data analysis methods, devised for moderate or small datasets, cope with the complexities of modern data? The case of high dimensional data is particularly revealing of some of the drawbacks. We look at the case where the number of characteristics measured in an object is at least the number of observed objects and conclude that this configuration leads to geometrical and mathematical oddities and is an insurmountable barrier for the direct application of traditional methodologies. If scientists are going to ignore fundamental mathematical results arrived at in this paper and blindly use software to analyze data, the results of their analyses may not be trustful, and the findings of their experiments may never be validated. That is why new methods together with the wise use of traditional approaches are essential to progress safely through the present reality.

研究动机与目标

  • 揭示在 $ p \geq n $ 的高维数据集中应用经典多元数据分析方法时存在的数学与几何缺陷。
  • 证明在 $ p \geq n-1 $ 的高维空间中,任意两点间的马氏距离均趋于一致且退化。
  • 证明此类高维数据的任意二维投影均可被构造为与任意 $ n $ 个点的配置完全相似,无论原始数据结构如何。
  • 提醒研究人员避免盲目使用标准软件工具处理高维数据,因为结果可能在数学上无效且不可验证。
  • 倡导建立尊重高维数据内在几何特性的新分析框架,尤其在 $ p \geq n $ 的情况下。

提出的方法

  • 使用帽子矩阵 $ H = X(X^TX)^{-1}X^T $ 分析高维回归中的杠杆与影响,表明当 $ p \geq n-1 $ 时 $ H = I_n $。
  • 应用马氏距离公式 $ d^2(x_i, x_j) = (x_i - x_j)^T S^{-1} (x_i - x_j) $,并在 $ p \geq n-1 $ 的约束下推导其边界。
  • 建立理论边界:对所有 $ i \neq j $,有 $ d_{x_i, \bar{x}} \leq (n-1)n^{-1/2} $ 与 $ d_{x_i, x_j} \leq \{2(n-1)\}^{1/2} $,表明距离的退化性。
  • 通过奇异值分解(SVD)进行正交投影,构造变换 $ Q $,使得 $ XQ = Y^* $,其中 $ Y^* $ 是任意目标二维配置 $ Y $ 的仿射变换。
  • 使用标准化数据矩阵 $ Z = X_c S^{-1/2} $,并构造 $ U = Z^T Y / (n-1) $,证明当 $ Y $ 标准化时 $ U $ 具有正交列。
  • 证明当 $ p < n-1 $ 时,此类二维投影的完全相似性不可能实现,从而在 $ p = n-1 $ 处确立了明确的临界转变。

实验结果

研究问题

  • RQ1当 $ p \geq n-1 $ 时,数据点之间的马氏距离会发生什么变化?这对经典多元推断有何影响?
  • RQ2能否将任意 $ n $ 个点的二维配置精确重现为 $ p \geq n-1 $ 的高维数据集的投影?
  • RQ3为何传统数据分析方法在 $ p \geq n $ 的高维环境中失效?其背后的数学原理是什么?
  • RQ4当 $ p \geq n-1 $ 时,帽子矩阵 $ H $ 的行为如何?这对回归中杠杆与影响的诊断有何影响?
  • RQ5为确保在 $ p \geq n $ 时分析的有效性,必须施加哪些约束?如何判断此类方法已不再适用?

主要发现

  • 当 $ p \geq n-1 $ 时,任意两点间马氏距离的上界为 $ \{2(n-1)\}^{1/2} $,表明高维空间中点间有意义的分离性丧失。
  • 每个数据点到样本均值的距离上界为 $ (n-1)n^{-1/2} $,表明随着 $ n $ 增大,所有点均位于一个不断收缩的球壳内。
  • 当 $ p \geq n-1 $ 时,帽子矩阵变为 $ H = I_n $,意味着所有杠杆值 $ h_{ii} = 1 $,从而导致标准影响诊断方法失效。
  • 对于 $ p \geq n-1 $ 的高维数据集,其任意二维投影均可精确匹配(至仿射变换)任意 $ n $ 个点的配置,无论原始数据结构如何。
  • 此类完美投影的存在意味着当 $ p \geq n-1 $ 时,高维数据的二维可视化本质上不可靠,因为它们可能模仿任意任意模式。
  • 当 $ p < n-1 $ 时,此类二维投影的完全相似性不可能实现,表明在 $ p = n-1 $ 处存在明确的几何转变,该点定义了数据分析有效性的临界阈值。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。