[論文レビュー] Is Extreme Learning Machine Feasible? A Theoretical Assessment (Part II)
本稿は、極端な学習機械(ELM)の理論的妥当性を評価し、隠層の重みをランダム化することで計算コストを削減できるものの、2つの主要な欠陥を示している。 (1) ランダム性による近似と学習の内在的不確実性、(2) ガウス型活性化関数を用いた場合の一般化性能の低下である。研究では、ガウスカーネルを用いたELMが標準的なFNNより性能が劣ることを証明しているが、係数正則化と複数回の学習試行によりこれを是正でき、期待される最適な学習速度が回復することを示している。
An extreme learning machine (ELM) can be regarded as a two stage feed-forward neural network (FNN) learning system which randomly assigns the connections with and within hidden neurons in the first stage and tunes the connections with output neurons in the second stage. Therefore, ELM training is essentially a linear learning problem, which significantly reduces the computational burden. Numerous applications show that such a computation burden reduction does not degrade the generalization capability. It has, however, been open that whether this is true in theory. The aim of our work is to study the theoretical feasibility of ELM by analyzing the pros and cons of ELM. In the previous part on this topic, we pointed out that via appropriate selection of the activation function, ELM does not degrade the generalization capability in the expectation sense. In this paper, we launch the study in a different direction and show that the randomness of ELM also leads to certain negative consequences. On one hand, we find that the randomness causes an additional uncertainty problem of ELM, both in approximation and learning. On the other hand, we theoretically justify that there also exists an activation function such that the corresponding ELM degrades the generalization capability. In particular, we prove that the generalization capability of ELM with Gaussian kernel is essentially worse than that of FNN with Gaussian kernel. To facilitate the use of ELM, we also provide a remedy to such a degradation. We find that the well-developed coefficient regularization technique can essentially improve the generalization capability. The obtained results reveal the essential characteristic of ELM and give theoretical guidance concerning how to use ELM.
研究の動機と目的
- ELMの実証的成功を超えて、理論的妥当性を評価すること。
- ELMの隠層における重みのランダム割り当てがもたらす悪影響を調査すること。
- 特定の活性化関数を用いた場合にELMの一般化能力が低下するかどうかを特定すること。
- ELMにおける不確実性と一般化能力の低下を是正する手法を提案すること。
- ELMにおける活性化関数の選択と正則化戦略の理論的基盤を確立すること。
提案手法
- 統計的学習理論の枠組み内でELMを分析し、一般化誤差と近似誤差に注目する。
- 集中不等式と正則化技術を用いて期待一般化誤差の上限を導出する。
- $l^2$係数正則化を適用して出力重みの学習を安定化させ、一般化性能を向上させる。
- 複数回の学習試行を用いて、重み初期化のランダム性がもたらす不確実性を軽減する。
- 標本依存の仮説空間とカーネルベースの学習モデルを用いて理論的誤差上限を導出する。
- ガウスカーネルを用いたELMと標準的なFNNを比較し、正則化なしではELMの収束速度が最適でないことを示す。
実験結果
リサーチクエスチョン
- RQ1ELMの隠層重みにおけるランダム性が、近似および学習の両方において不確実性を引き起こすか?
- RQ2ランダム重み割り当てであっても、特定の活性化関数を用いた場合にELMの一般化能力が低下するか?
- RQ3ガウス型活性化関数を用いたELMの一般化性能は、標準的なFNNより劣るのか?
- RQ4$l^2$正則化により、特にガウスカーネルを用いた場合にELMが最適な学習速度を回復できるか?
- RQ5ELMの一般化能力が低下するか否かを決定する活性化関数の条件は何か?
主な発見
- 理論的比較により、ガウス型活性化関数を用いたELMは、標準的なFNNと比較して一般化能力が劣ることが示された。
- 同じ条件下で、ガウスカーネルを用いたELMの一般化誤差は、FNNよりも根本的に悪い。
- 複数回の学習試行は、ELMにおける不確実性問題を効果的に軽減し、単一実行における失敗リスクを低減する。
- $l^2$係数正則化を出力重みに適用することで、ELMは期待される観点でほぼ最適な学習速度を達成でき、FNNの性能と一致する。
- 正則化されたELMの理論的誤差上限は、$O(m^{-\frac{2r}{2r+d} + \varepsilon})$ のオーダーに従い、正則化が適切に調整されればFNNの最適レートに近づく。
- ELMの一般化能力に与える影響に基づく活性化関数の分類基準は、一般に未解決の問題であるが、本稿は今後の研究の基盤を提供している。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。