Skip to main content
QUICK REVIEW

[論文レビュー] Refined Complexity of PCA with Outliers

Fedor V. Fomin, Petr A. Golovach|arXiv (Cornell University)|May 10, 2019
Sparse and Compressive Sensing Techniques被引用数 6
ひとこと要約

この論文は、外れ値を含むPCAのための固定パラメータ可 tractable(FPT)なアルゴリズムを提示しており、時間計算量 $ n^{/mathcal{O}(d^2)} $ で問題を解ける。これは定数次元 $ d $ に対して多項式時間である。また、指数時間仮説(ETH)が成り立たない限り、任意の関数 $ f $ に対して、時間 $ f(d)n^{o(d)} $ で問題を解くこともしくは定数因子で近似することは不可能であり、問題のタイトな計算量下界を確立している。

ABSTRACT

Principal component analysis (PCA) is one of the most fundamental procedures in exploratory data analysis and is the basic step in applications ranging from quantitative finance and bioinformatics to image analysis and neuroscience. However, it is well-documented that the applicability of PCA in many real scenarios could be constrained by an "immune deficiency" to outliers such as corrupted observations. We consider the following algorithmic question about the PCA with outliers. For a set of $n$ points in $\mathbb{R}^{d}$, how to learn a subset of points, say 1% of the total number of points, such that the remaining part of the points is best fit into some unknown $r$-dimensional subspace? We provide a rigorous algorithmic analysis of the problem. We show that the problem is solvable in time $n^{O(d^2)}$. In particular, for constant dimension the problem is solvable in polynomial time. We complement the algorithmic result by the lower bound, showing that unless Exponential Time Hypothesis fails, in time $f(d)n^{o(d)}$, for any function $f$ of $d$, it is impossible not only to solve the problem exactly but even to approximate it within a constant factor.

研究の動機と目的

  • 実世界のデータ解析における古典的PCAの外れ値に対する脆弱性に対処すること。
  • 外れ値を同定し除外することを目的とするPCAのパラメータ化された複雑性を形式的に研究すること。
  • 次元 $ d $ および外れ値数 $ k $ の観点から、問題の正確な計算量複雑性を特定すること。
  • 標準的な複雑性仮定の下で、問題の tractability に関するタイトな上界および下界を確立すること。

提案手法

  • 外れ値付きPCAを、スパースな不正を伴う低ランク近似問題として定式化:$ \|M - L - S\|_F^2 $ を最小化し、ここで $ \operatorname{rank}(L) \leq r $ かつ $ S $ は高々 $ k $ 個の非ゼロ行を持つ。
  • 時間計算量 $ n^{\mathcal{O}(d^2)} $ で動作するアルゴリズムを開発し、$ d $ が定数のとき多項式時間で動作することを示した。
  • 問題の複雑性に下界を示すために、マルチカラーモノクロームクライQUE問題への還元を採用した。
  • 行列摂動およびランク解析を用いた代数的・幾何的議論により、小さな誤差を持つ解がマルチカラーモノクロームクライQUEの解を意味することを示した。
  • 指数時間仮説(ETH)を用いて、任意の関数 $ f $ に対して、時間 $ f(d)n^{o(d)} $ で動作するアルゴリズムは、定数因子近似でさえも除外できる。
  • 外れ値数 $ k $ にパラメータ化した場合、問題は $ \operatorname{{\sf W}}[1] $-難易度であることが示され、$ \operatorname{{\sf FPT}} = \operatorname{{\sf W}}[1] $ でない限り、$ f(d)N^{o(d)} $ 時間で動作する $ \omega $-近似アルゴリズムは存在しない。

実験結果

リサーチクエスチョン

  • RQ1次元 $ d $ が小さい場合、一般にNP困難であるにもかかわらず、PCAと外れ値は効率的に解けるか?
  • RQ2次元 $ d $ にパラメータ化した場合のPCAと外れ値の正確なパラメータ化された複雑性は何か?
  • RQ3次元 $ d $ に関して、時間 $ d $ の部分指数的時間で動作する定数因子近似アルゴリズムは存在するか?
  • RQ4指数時間仮説を仮定した場合、任意の関数 $ f $ に対して、時間 $ f(d)n^{o(d)} $ で問題を解けるか?
  • RQ5PCAと外れ値の複雑性とマルチカラーモノクロームクライQUE問題との関係は何か?

主な発見

  • PCAと外れ値の問題は、時間 $ n^{\mathcal{O}(d^2)} $ で解けるため、次元 $ d $ に関して固定パラメータ可 tractable(FPT)である。
  • 定数 $ d $ の場合、問題は多項式時間で解けるため、外れ値を含む低次元データに対して実用的な解決策を提供する。
  • 指数時間仮説(ETH)の下では、任意の関数 $ f $ に対して、時間 $ f(d)n^{o(d)} $ で問題を解くこともしくは定数因子で近似することは不可能である。
  • 外れ値数 $ k $ にパラメータ化した場合、問題は $ \operatorname{{\sf W}}[1] $-難易度であるため、このパラメータでは強い非可 tractability が示唆される。
  • 任意の計算可能関数 $ f $ に対して、時間 $ f(d)N^{o(d)} $ で動作する $ \omega $-近似アルゴリズムは存在しない($ \operatorname{{\sf FPT}} = \operatorname{{\sf W}}[1] $ でない限り)、ここで $ N $ は行列のビットサイズである。
  • 下界の結果はタイトであり、上界と一致しており、$ n^{\mathcal{O}(d^2)} $ 時間のアルゴリズムがETHの下で本質的に最適であることが示された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。