Skip to main content
QUICK REVIEW

[論文レビュー] Optical Transformers

Maxwell G. Anderson, Shi-Yuan Ma|arXiv (Cornell University)|Feb 20, 2023
Neural Networks and Reservoir Computing被引用数 6
ひとこと要約

本稿では、光学的行列-ベクタ乗算を用いることで、デジタルシステムに比べて著しく高いエネルギー効率を達成する神経ネットワークアクセラレータ「光学的トランスフォーマー」を提案する。実験的な光学ハードウェア上で正確な推論を実証し、大規模モデルのシミュレーションを実施した結果、光学システムは1/dのエネルギー/MACあたりの利点を示し、トランジスタ数が1兆を超えるモデルでは、最先端のデジタルプロセッサに比べ最大8,000倍のエネルギー効率向上が可能であることが示された。

ABSTRACT

The rapidly increasing size of deep-learning models has caused renewed and growing interest in alternatives to digital computers to dramatically reduce the energy cost of running state-of-the-art neural networks. Optical matrix-vector multipliers are best suited to performing computations with very large operands, which suggests that large Transformer models could be a good target for optical computing. To test this idea, we performed small-scale optical experiments with a prototype accelerator to demonstrate that Transformer operations can run on optical hardware despite noise and errors. Using simulations, validated by our experiments, we then explored the energy efficiency of optical implementations of Transformers and identified scaling laws for model performance with respect to optical energy usage. We found that the optical energy per multiply-accumulate (MAC) scales as $\frac{1}{d}$ where $d$ is the Transformer width, an asymptotic advantage over digital systems. We conclude that with well-engineered, large-scale optical hardware, it may be possible to achieve a $100 imes$ energy-efficiency advantage for running some of the largest current Transformer models, and that if both the models and the optical hardware are scaled to the quadrillion-parameter regime, optical computers could have a $>8,000 imes$ energy-efficiency advantage over state-of-the-art digital-electronic processors that achieve 300 fJ/MAC. We analyzed how these results motivate and inform the construction of future optical accelerators along with optics-amenable deep-learning approaches. With assumptions about future improvements to electronics and Transformer quantization techniques (5$ imes$ cheaper memory access, double the digital--analog conversion efficiency, and 4-bit precision), we estimated that optical computers' advantage against current 300-fJ/MAC digital processors could grow to $>100,000 imes$.

研究の動機と目的

  • 大規模なディープラーニングモデル、特にトランスフォーマーの学習および推論に伴うエネルギーコストの増大に対処すること。
  • 光学コンピューティングが、大規模なトランスフォーマー・モデル向けにスケーラブルかつエネルギー効率の高いデジタル電子プロセッサの代替手段となり得るかどうかを調査すること。
  • ノイズや誤差が存在する中でも、光学ハードウェアがトランスフォーマー演算を正確に実行できることを実証すること。
  • 異なるモデルサイズおよび光学的エネルギー予算における光学的エネルギー効率とモデル性能のスケーリング則を確立すること。
  • 実験的およびシミュレーテッドな結果に基づき、将来の光学アクセラレータおよび光学に適したディープラーニングアーキテクチャの設計を支援すること。

提案手法

  • トランスフォーマー推論に代表される行列-ベクタ乗算演算を実行するため、空間光変調器(SLM)を用いたプロトタイプを用いて小規模な光学実験を実施した。
  • 光学ハードウェアから得た実世界のノイズ、誤差、不正確さのデータを収集し、完全な光学的トランスフォーマーの高精度なシミュレーションをキャリブレーションした。
  • 実験的測定から得た系統的な誤差、ノイズ、重み/入力の不正確さを用いて、光学的トランスフォーマー推論のシミュレーションを実施した。
  • 光学的エネルギー/MAC、メモリアクセス、およびデジタル-アナログ変換のオーバーヘッドを含めた光学ニューラルネットワーク(ONN)アクセラレータの総エネルギー消費量をモデル化した。
  • モデル幅(d)および総光学的エネルギー使用量を関数としてエネルギー効率のスケーリング則を評価し、光学的エネルギー/MACが1/dのスケーリングを示すことを明らかにした。
  • より良い電子回路、メモリアクセス、および4ビット量子化を想定した条件下で、性能およびエネルギー利点を予測した。
Figure 1: General scheme of an optical neural network (ONN) accelerator. Data is encoded and fed into the network, and the output is subject to shot noise. There are many experimental realizations of ONN accelerators such as Mach-Zehnder Interferometer meshes (Shen et al., 2017 ; Bogaerts et al., 20
Figure 1: General scheme of an optical neural network (ONN) accelerator. Data is encoded and fed into the network, and the output is subject to shot noise. There are many experimental realizations of ONN accelerators such as Mach-Zehnder Interferometer meshes (Shen et al., 2017 ; Bogaerts et al., 20

実験結果

リサーチクエスチョン

  • RQ1ノイズや誤差が内在する中でも、光学ハードウェアはトランスフォーマーに必要な線形演算を正確に実行できるか?
  • RQ2モデルサイズおよび光学的エネルギー使用量の変化に伴い、光学的トランスフォーマーのエネルギー効率はどのように変化するか?
  • RQ3大規模なトランスフォーマー推論において、光学プロセッサがデジタルプロセッサに比べて理論的および実用的なエネルギー効率の優位性を示すか?
  • RQ4学習および量子化スキームは、光学的トランスフォーマーにおける光子使用量と性能にどのように影響を与えるか?
  • RQ5将来の1兆トランジスタ規模のモデルをターゲットとする光学アクセラレータの設計における主要なトレードオフとスケーリング限界は何か?

主な発見

  • 実機でのノイズや誤差が存在する中でも、光学的行列-ベクタ乗算は高い正確性で実行可能であり、実験およびシミュレーションによって検証された。
  • 光学的エネルギー/乗算-積算(MAC)演算は、モデル幅dの関数として1/dのスケーリングを示し、エネルギー/MACあたりの定常的エネルギー優位性をデジタルシステムが提供する。
  • 現在の大型トランスフォーマー(例:175Bパラメータモデル)では、光学システムが最先端のデジタルプロセッサ(300 fJ/MAC)に比べ100倍のエネルギー効率向上を達成できる。
  • モデルおよび光学ハードウェアを1京パラメータ規模にスケーリングした場合、光学システムはデジタルプロセッサに比べ8,000倍以上のエネルギー効率向上を達成できる。
  • 楽観的な仮定(メモリアクセスが5倍安価、デジタル-アナログ変換効率が2倍向上、4ビット量子化)のもとでは、エネルギー優位性が100,000倍を超える可能性がある。
  • 8ビットのデジタルモデル性能を維持するための光子使用量は、モデルサイズに従って非線形的に増加するが、重みおよび活性化の統計に強く依存しており、これは量子化および学習スキームに影響を受ける。
Figure 2: Optical Transformer evaluation: prototype hardware; simulator model; Transformer architecture. Bottom: typical Transformer architecture, but with ReLU6 activation. Top Left: experimental spatial light modulator (SLM)-based accelerator setup. From some layers—marked with a laser icon—we sam
Figure 2: Optical Transformer evaluation: prototype hardware; simulator model; Transformer architecture. Bottom: typical Transformer architecture, but with ReLU6 activation. Top Left: experimental spatial light modulator (SLM)-based accelerator setup. From some layers—marked with a laser icon—we sam

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。