Skip to main content
QUICK REVIEW

[論文レビュー] A locality-based approach for coded computation

Michael Rudow, K. V. Rashmi|arXiv (Cornell University)|Feb 6, 2020
Advanced Data Storage Technologies参考文献 54被引用数 9
ひとこと要約

本論文は、誤り訂正符号の局所性特性を計算耐障害性にモデル化する、局所性に基づく符号化計算フレームワークを導入し、多次元多項式計算において、従来の手法よりも少ないワーカー数を実現する。リード・ムラー符号の局所的回復方式を活用することで、計算局所性を伴いながらs人の遅延ワーカーに対しても耐障害性を達成し、非線形関数に対しては従来のs倍のワーカー過剰負荷の壁を破る。

ABSTRACT

Modern distributed computation infrastructures are often plagued by unavailabilities such as failing or slow servers. These unavailabilities adversely affect the tail latency of computation in distributed infrastructures. The simple solution of replicating computation entails significant resource overhead. Coded computation has emerged as a resource-efficient alternative, wherein multiple units of data are encoded to create parity units and the function to be computed is applied to each of these units on distinct servers. A decoder can use the available function outputs to decode the unavailable ones. Existing coded computation approaches are resource efficient only for simple variants of linear functions such as multilinear, with even the class of low degree polynomials requiring the same multiplicative overhead as replication for practically relevant straggler tolerance. In this paper, we present a new approach to model coded computation via the lens of locality of codes. We introduce a generalized notion of locality, denoted computational locality, building upon the locality of an appropriately defined code. We show that computational locality is equivalent to the required number of workers for coded computation and leverage results from the well-studied locality of codes to design coded computation schemes. We show that recent results on coded computation of multivariate polynomials can be derived using local recovering schemes for Reed-Muller codes. We present coded computation schemes for multivariate polynomials that adaptively exploit locality properties of input data-- an inadmissible technique under existing frameworks. These schemes require fewer workers than the lower bound under existing coded computation frameworks, showing that the existing multiplicative overhead on the number of servers is not fundamental for coded computation of nonlinear functions.

研究の動機と目的

  • 非線形関数、特に多次元多項式に対して、従来の符号化計算手法の高いリソース過剰負荷を解消すること。
  • 従来の手法が、たとえ低次の非線形関数であっても、s人の遅延ワーカーを耐えられるようにするにはs倍のワーカーを必要としているという根本的制限を克服すること。
  • 計算耐障害性と下位の符号の局所性を結びつける新しい理論的枠組みを確立すること。
  • 入力データの局所性を適応的に活用する符号化計算スキームを設計し、従来の下限を下回るワーカー数を実現すること。
  • s倍のワーカー過剰負荷が、非線形符号化計算において根本的な限界ではないことを示すこと。

提案手法

  • 符号の局所性に直接関連する、計算局所性の新しい概念を導入し、これにより必要なワーカー数に結びつける。
  • 各関数評価を符号語に対応させた符号として符号化計算をモデル化し、パuncturedコードを用いて局所的回復を分析する。
  • リード・ムラー符号の既知の局所的回復方式を応用し、多次元多項式のための符号化計算スキームを構築する。
  • データ固有の局所性を活用する適応的スキームを設計し、従来の下限を下回るワーカー数を実現する。
  • 多項式合成の補間を用いて、部分的なワーカー結果から出力を回復する。計算局所性がワーカー数を決定する。
  • 既存のスキーム(多項式符号やMatDot符号)を、符号の局所性の観点から再解釈し、それらの下位構造を明らかにする。

実験結果

リサーチクエスチョン

  • RQ1非線形関数の符号化計算におけるワーカー過剰負荷を、従来の手法が要求するs倍要因未満に低減できるか。
  • RQ2s倍のワーカー過剰負荷は、多次元多項式の符号化計算において根本的な限界であるか。
  • RQ3符号の局所性特性を体系的に活用して、より効率的な符号化計算スキームを設計できるか。
  • RQ4従来の入力に依存しないフレームワークでは不適切な入力固有のデータ局所性を、どのように活用できるか。
  • RQ5よく研究された符号の局所性の結果を、符号化計算分野にどの程度まで転用できるか。

主な発見

  • 提案された局所性ベースのモデルにより、符号化計算における計算局所性と必要な最小ワーカー数の等価性が確立される。
  • 本手法により、従来のフレームワークの下限を下回るワーカー数で、多次元多項式の符号化計算スキームを実現でき、s倍の過剰負荷の壁を破る。
  • 次数(s+1)の多次元多項式に対しては、s+1人のワーカーでs人の遅延ワーカーを耐えられ、従来の手法がs×(s+1)人を必要としていたのに対し、著しく効率的である。
  • 本手法により、リード・ムラー符号の局所性に基づく統一的フレームワークを通じて、既存の最適スキーム(多項式符号やMatDot符号)を再現できる。
  • MatDot符号における符号記号の計算局所性は2t−1+sで抑えられるが、これは多項式符号のt²+sの境界よりも顕著に低い。通信量の増加を犠牲にしているが、その代わりに高い効率性を達成する。
  • 本フレームワークにより、入力データの局所性を活用する適応的スキームが可能となり、従来の入力に依存しないモデルでは実現できない。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。