Skip to main content
QUICK REVIEW

[論文レビュー] Adaptive Importance Sampling for Finite-Sum Optimization and Sampling with Decreasing Step-Sizes

Ayoub El Hanchi, David A. Stephens|arXiv (Cornell University)|Mar 23, 2021
Stochastic Gradient Optimization Techniques被引用数 4
ひとこと要約

本稿では、減少するステップサイズを用いた有限和最適化およびサンプリングのための適応的 Importance Sampling アルゴリズムである Avare を提案する。オンライン学習と分散低減技術を活用することで、SGD と SGLD に対してそれぞれ $Ø(T^{2/3})$ および $Ø(T^{5/6})$ の動的リグレットを達成し、分散が支配的になる後段階の最適化で収束性が著しく向上する。

ABSTRACT

Reducing the variance of the gradient estimator is known to improve the convergence rate of stochastic gradient-based optimization and sampling algorithms. One way of achieving variance reduction is to design importance sampling strategies. Recently, the problem of designing such schemes was formulated as an online learning problem with bandit feedback, and algorithms with sub-linear static regret were designed. In this work, we build on this framework and propose Avare, a simple and efficient algorithm for adaptive importance sampling for finite-sum optimization and sampling with decreasing step-sizes. Under standard technical conditions, we show that Avare achieves $\mathcal{O}(T^{2/3})$ and $\mathcal{O}(T^{5/6})$ dynamic regret for SGD and SGLD respectively when run with $\mathcal{O}(1/t)$ step sizes. We achieve this dynamic regret bound by leveraging our knowledge of the dynamics defined by the algorithm, and combining ideas from online learning and variance-reduced stochastic optimization. We validate empirically the performance of our algorithm and identify settings in which it leads to significant improvements.

研究の動機と目的

  • 大規模な有限和問題における確率的勾配推定の分散の課題に対処すること。
  • 減少するステップサイズを伴う最適化中に、分散を動的に最小化する適応的 Importance Sampling 策略の開発。
  • 標準的な技術的条件下で、SGD および SGLD の両方に対して非線形の動的リグレット境界を達成すること。
  • 実用的な機械学習設定における提案手法の性能向上を実証的に検証すること。
  • アンサンブルサンプリングをミニバッチ形式に拡張しつつ、不偏性と分散低減を維持すること。

提案手法

  • Avare は、部分的勾配ノルムの観測を用いて、バンドイットフィードバック付きのオンライン学習問題として Importance Sampling を定式化する。
  • 勾配推定器の共分散行列のトレースを最小化するように、動的にスケーリング分布 $p^t$ を更新する。
  • サンプリングなしで置換する手法と Importance Sampling の利点を組み合わせた、新規のミニバッチ推定器を採用し、不偏性を維持する。
  • ステップサイズの減少といったアルゴリズムのダイナミクスの知識を活用して、リグレット境界を導出する。
  • 定常ステップサイズへの適応を可能にするために、修正されたエプシロン列を組み込み、安定した性能を確保する。
  • オンライン学習と分散低減付き確率的最適化のツールを組み合わせ、動的リグレット境界を導出する。

実験結果

リサーチクエスチョン

  • RQ1適応的 Importance Sampling は、減少するステップサイズを用いた SGD および SGLD に対して、非線形の動的リグレットを達成できるか?
  • RQ2非定常設定下で、適応的 Importance Sampling の動的リグレットは、静的リグレットと比べてどのように異なるか?
  • RQ3ステップサイズの減少は、確率的最適化における分散低減の有効性にどのような影響を与えるか?
  • RQ4サンプリングなしで置換する手法と Importance Sampling の両方の利点を組み合わせつつ、不偏性を維持できるミニバッチ推定器を設計できるか?
  • RQ5実世界のデータセットにおいて、Avare は一様サンプリングや他の Importance Sampling 策略と比べて、どのように実験的に性能を発揮するか?

主な発見

  • 標準的な技術的条件下で、$Ø(1/t)$ ステップサイズを用いる場合、Avare は SGD に対して $Ø(T^{2/3})$ の動的リグレットを達成する。
  • 同じステップサイズの設定下で、SGLD に対して Avare は $Ø(T^{5/6})$ の動的リグレットを達成する。
  • アルゴリズムは勾配推定器の分散を著しく低減し、最適化の後段階での収束を早める。
  • MNIST、IJCNN1、CIFAR10 における実験結果から、Avare は一様サンプリングや他のベースラインを上回り、特に分散支配的領域で顕著な性能向上を示す。
  • 修正されたエプシロン列により、定常ステップサイズでも安定した性能が得られ、減少するステップサイズの仮定を超えたロバストネスを示唆する。
  • 提案されたミニバッチ推定器は、両方の利点を組み合わせつつ、不偏性を維持する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。