[論文レビュー] A Survey of Big Data Machine Learning Applications Optimization in Cloud Data Centers and Networks
本サーベイは、クラウドデータセンターおよびネットワークにおけるビッグデータ機械学習ワークロードの最適化技術について包括的な分析を提供し、アプリケーションレベル、ネットワーキングレベル、データセンターレベルの最適化に分類している。主な課題として、トラフィックの混雑、エネルギー消費、マルチテナント環境における非効率性を特定し、仮想化、SDN、NFV、コンテナ化を活用したソリューションを評価することで、分散システム全体におけるパフォーマンス、公平性、エネルギー効率の向上を検討している。
This survey article reviews the challenges associated with deploying and optimizing big data applications and machine learning algorithms in cloud data centers and networks. The MapReduce programming model and its widely-used open-source platform; Hadoop, are enabling the development of a large number of cloud-based services and big data applications. MapReduce and Hadoop thus introduce innovative, efficient, and accelerated intensive computations and analytics. These services usually utilize commodity clusters within geographically-distributed data centers and provide cost-effective and elastic solutions. However, the increasing traffic between and within the data centers that migrate, store, and process big data, is becoming a bottleneck that calls for enhanced infrastructures capable of reducing the congestion and power consumption. Moreover, enterprises with multiple tenants requesting various big data services are challenged by the need to optimize leasing their resources at reduced running costs and power consumption while avoiding under or over utilization. In this survey, we present a summary of the characteristics of various big data programming models and applications and provide a review of cloud computing infrastructures, and related technologies such as virtualization, and software-defined networking that increasingly support big data systems. Moreover, we provide a brief review of data centers topologies, routing protocols, and traffic characteristics, and emphasize the implications of big data on such cloud data centers and their supporting networks. Wide ranging efforts were devoted to optimize systems that handle big data in terms of various applications performance metrics and/or infrastructure energy efficiency. Finally, some insights and future research directions are provided.
研究の動機と目的
- クラウド環境におけるビッグデータ機械学習アプリケーションの最適化戦略を特定および分類すること。
- トラフィックの混雑、エネルギー消費、リソースの過小/過剰利用といった、クラウドデータセンターおよびネットワークにおける課題を分析すること。
- SDN、NFV、コンテナなど、新興技術がビッグデータシステムのパフォーマンスおよび効率に与える影響を評価すること。
- アプリケーション、ネットワーキング、データセンターの各レイヤーにわたる最適化技術を体系的にレビューすること。
- スケーラブルでエネルギー効率が高く、高性能なビッグデータシステムを実現するための研究ギャップおよび今後の方向性を強調すること。
提案手法
- クラウドデータセンターおよびネットワークにおけるビッグデータおよび機械学習最適化に関する既存文献の体系的レビュー。
- 最適化研究を3つのカテゴリーに分類:アプリケーションレベル最適化、ネットワーキングレベル最適化、データセンターレベル最適化。
- MapReduce、Hadoop、SDN、NFV、仮想マシン、コンテナ、データセンターのトポロジーを含む、主要な技術の分析。
- 完了時間、公平性、コスト、利益、エネルギー消費といったパフォーマンス指標を、多様なワークロードで評価。
- シミュレーション結果および実世界のプロトタイプやクラウドテストベッドからの実験結果の統合。
- 分散およびグローバル分散フレームワークにおけるパフォーマンス、エネルギー効率、リソース利用率のトレードオフの特定。
実験結果
リサーチクエスチョン
- RQ1ビッグデータワークロードは、トラフィック、遅延、リソース利用率の観点から、クラウドデータセンターおよびネットワークのパフォーマンスにどのように影響を与えるか?
- RQ2クラウドスタックの異なるレイヤー(アプリケーション、ネットワーク、データセンター)において、ビッグデータ機械学習アプリケーションの最適化に直面する主な課題は何か?
- RQ3SDN、NFV、コンテナといった新興技術は、どのようにビッグデータシステムのエネルギー効率およびパフォーマンスを向上させることができるか?
- RQ4マルチテナントおよびグローバル分散環境におけるビッグデータ環境において、パフォーマンス、コスト、エネルギー消費のトレードオフは何か?
- RQ5クラウドインfraストラクチャにおいて、スケーラブルで公平かつエネルギー効率の高いビッグデータ処理を実現するための未解決の研究課題は何か?
主な発見
- エネルギー効率とパフォーマンスはしばしば相反し、多くのプロバイダーがSLAを満たすために過剰なリソースを割り当てることでエネルギー消費を最小限に抑えるのではなく、優先している。
- SDNおよびNFVは、動的でアプリケーションに適応したネットワークおよびリソース管理を可能にし、ジョブ完了時間を短縮するとともにエネルギー効率を向上させている。
- マルチテナント環境では、共有されたネットワークおよびI/Oリソースのため、公平性と隔離性の課題が生じ、動的スケジューリングおよび価格設定モデルの導入が不可欠である。
- グローバル分散フレームワークでは、高い遅延とデータ転送コストが課題となり、新たなルーティングおよびリソース割り当て戦略の導入が不可欠である。
- 異種のクラスタはタスク完了時間の不均衡を引き起こし、正確なプロファイリングおよび知能的なスケジューリングアルゴリズムの導入が求められる。
- コンテナ化および仮想化は、特に動的かつ大規模なデータセンター環境において、リソース利用率の向上と柔軟性の向上を実現している。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。