[论文解读] A Survey on Deep Neural Network Partition over Cloud, Edge and End Devices
本综述对云、边缘和终端设备上的深度神经网络(DNN)划分进行了全面分析,提出了一套五维分类框架,统一了现有方法。该研究构建了标准化的DNN划分策略评估模型,识别出动态环境、隐私保护和算法复杂性等关键挑战,为分布式AI推理的未来研究奠定了基础。
Deep neural network (DNN) partition is a research problem that involves splitting a DNN into multiple parts and offloading them to specific locations. Because of the recent advancement in multi-access edge computing and edge intelligence, DNN partition has been considered as a powerful tool for improving DNN inference performance when the computing resources of edge and end devices are limited and the remote transmission of data from these devices to clouds is costly. This paper provides a comprehensive survey on the recent advances and challenges in DNN partition approaches over the cloud, edge, and end devices based on a detailed literature collection. We review how DNN partition works in various application scenarios, and provide a unified mathematical model of the DNN partition problem. We developed a five-dimensional classification framework for DNN partition approaches, consisting of deployment locations, partition granularity, partition constraints, optimization objectives, and optimization algorithms. Each existing DNN partition approache can be perfectly defined in this framework by instantiating each dimension into specific values. In addition, we suggest a set of metrics for comparing and evaluating the DNN partition approaches. Based on this, we identify and discuss research challenges that have not yet been investigated or fully addressed. We hope that this work helps DNN partition researchers by highlighting significant future research directions in this domain.
研究动机与目标
- 为解决在资源受限的终端和边缘设备上部署大规模DNN的挑战,通过在异构计算环境中实现高效的模型划分。
- 将文献中多样化的DNN划分方法统一到一个连贯且可扩展的分类框架中,以实现系统化的比较与分析。
- 识别并突出当前研究中尚未充分探索的关键挑战,包括动态环境、隐私保护和算法效率。
- 提出标准化的评估指标,用于在延迟、能耗和资源利用率等不同优化目标下,对DNN划分策略进行评估与比较。
- 通过提出动态、隐私感知和可扩展的DNN划分在边缘和终端设备推理中的有前景研究方向,为未来研究提供指导。
提出的方法
- 提出一种五维DNN划分分类框架,涵盖部署位置、划分粒度、约束条件、优化目标和优化算法。
- 构建DNN划分问题的统一数学模型,以形式化表达计算、通信和资源使用之间的权衡。
- 引入一组标准化评估指标,包括推理延迟、能耗、资源利用率和隐私影响,以实现不同方法之间的公平比较。
- 通过所提出的框架分析现有DNN划分技术,将每种方法映射到五个维度的具体取值上。
- 采用系统性文献综述方法,整合主要数字图书馆和搜索引擎中的研究成果,覆盖2015至2023年期间DNN划分的最新进展。
- 指出当前方法的局限性,例如忽略划分算法本身的计算复杂度,以及在卸载成本建模中采用静态假设。

实验结果
研究问题
- RQ1如何在不同部署环境(云、边缘、终端设备)中对DNN划分方法进行系统化分类与比较?
- RQ2DNN划分中的关键性能权衡是什么?如何在现实世界约束下对这些权衡进行建模与优化?
- RQ3当前DNN划分研究中的主要研究空白是什么,特别是在动态环境、隐私保护和算法效率方面?
- RQ4如何最小化DNN划分算法的计算复杂度,以支持大规模模型在实时应用中的部署?
- RQ5协同推理在提升分布式DNN推理中的隐私保护与性能方面发挥什么作用?
主要发现
- 五维分类框架成功地对现有DNN划分方法进行了映射与分类,实现了对多样化技术的系统性理解与比较。
- 当前DNN划分方法通常忽略划分算法本身的计算复杂度,导致在大规模模型上面临可扩展性问题。
- 在动态环境中(尤其是移动边缘设备)的卸载成本建模仍不充分,大多数研究假设网络条件静态或理想化。
- 资源利用率与公平分配是关键但被低估的优化目标,常被延迟和能耗等目标所掩盖。
- 隐私问题尤为突出,特别是在中间特征被传输至云端时,现有方法对隐私保护推理的支持有限。
- 由于移动性和设备可用性变化,动态DNN划分对现实世界物联网应用至关重要,但支持低开销运行时自适应的方法仍较少。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。