Skip to main content
QUICK REVIEW

[論文レビュー] Data-dependent compression of random features for large-scale kernel approximation

Raj Agrawal, Trevor Campbell|arXiv (Cornell University)|Oct 9, 2018
Stochastic Gradient Optimization Techniques被引用数 13
ひとこと要約

本稿では、理論的保証を維持したまま、大量のランダム特徴量を少数の重み付き部分集合に圧縮するデータ依存型圧縮手法を提案する。ランダム特徴量マップと効率的なデータ駆動型特徴選択を組み合わせることで、状態技術の手法よりもはるかに少ない特徴量 O(log J₊) を用いて、近似的に最適なカーネル行列近似を達成する。5000万件を超える観測値を含むデータセット上で実証された。

ABSTRACT

Kernel methods offer the flexibility to learn complex relationships in modern, large data sets while enjoying strong theoretical guarantees on quality. Unfortunately, these methods typically require cubic running time in the data set size, a prohibitive cost in the large-data setting. Random feature maps (RFMs) and the Nystrom method both consider low-rank approximations to the kernel matrix as a potential solution. But, in order to achieve desirable theoretical guarantees, the former may require a prohibitively large number of features J+, and the latter may be prohibitively expensive for high-dimensional problems. We propose to combine the simplicity and generality of RFMs with a data-dependent feature selection scheme to achieve desirable theoretical approximation properties of Nystrom with just O(log J+) features. Our key insight is to begin with a large set of random features, then reduce them to a small number of weighted features in a data-dependent, computationally efficient way, while preserving the statistical guarantees of using the original large set of features. We demonstrate the efficacy of our method with theory and experiments--including on a data set with over 50 million observations. In particular, we show that our method achieves small kernel matrix approximation error and better test set accuracy with provably fewer random features than state-of-the-art methods.

研究の動機と目的

  • 大規模な設定におけるカーネル手法の高い計算コストを軽減するため、正確なカーネル近似に必要なランダム特徴量の数を削減すること。
  • Nyström法の理論的頑健性とランダム特徴量マップの単純さと一般性を統合すること。
  • 元の大きな特徴集合の統計的保証を維持する、計算効率の良いデータ依存型特徴選択スキームを開発すること。
  • 既存手法よりも顕著に少ない特徴量で、証明可能なより良い近似誤差とテスト精度を達成すること。

提案手法

  • 理論的近似保証を確保するため、J₊ 個の大きなランダム特徴量のセットを出発点とする。
  • 初期セットから少数の重み付き特徴量部分集合を選択するデータ依存型圧縮スキームを適用する。
  • データの幾何構造とカーネル構造に基づいて、最も情報量の多い特徴量を特定するための凸最適化フレームワークを用いる。
  • 凸包と距離に基づく基準を活用し、データ多様体を最もよく表現する特徴量を優先順位付けする。
  • 圧縮された特徴量セットが、元のセットと同等の統計的性質を維持することを保証する。これには、カーネル行列近似誤差の境界も含まれる。
  • 5000万件を超える観測値を含むような巨大データセットにも効率的にスケーリングできるように、選択プロセスを最適化する。

実験結果

リサーチクエスチョン

  • RQ1理論的保証を維持したまま、カーネル近似に必要なランダム特徴量の数をデータ依存型圧縮スキームで削減できるか?
  • RQ2近似誤差とテスト精度の観点で、提案手法はNyström法および標準的ランダム特徴量マップと比べてどのように異なるか?
  • RQ3データ依存型選択のもとで、近似的に最適なカーネル近似を達成するために必要な最小特徴量は何か?
  • RQ45000万件を超える巨大データセットに対しても、計算コストを低く保ちながらスケーラブルか?

主な発見

  • 本手法は、完全なランダム特徴量マップと同等のカーネル行列近似誤差を達成するが、特徴量は O(log J₊) 個にまで圧縮され、特徴量数が顕著に削減されている。
  • 5000万件を超える観測値を含むデータセットにおいて、状態技術のアプローチよりもはるかに少ない特徴量で、高いテストセット精度を維持している。
  • データ依存型圧縮スキームは、元の大きな特徴集合の理論的近似保証を保持しており、頑健な性能を保証している。
  • 特徴次元が著しく削減された状態で、標準的ランダム特徴量マップおよびNyströmベースの手法よりも、近似誤差とテスト精度の両面で優れている。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。