Skip to main content
QUICK REVIEW

[论文解读] Deep Structured Teams with Linear Quadratic Model: Partial Equivariance and Gauge Transformation

Jalal Arabneydi, Amir G. Aghdam|arXiv (Cornell University)|Dec 9, 2019
Distributed Control Multi-Agent Systems参考文献 20被引用 9
一句话总结

本文提出具有线性二次动态的深度结构化团队,通过部分等变性和规范变换,推导出一种低复杂度、可扩展的大规模去中心化控制解决方案。在深度状态共享下,提出了闭式最优策略;在部分深度状态共享下,提出了次优策略,其计算复杂度与团队规模无关,并且随着代理数量的增加,信息和鲁棒性价格均收敛至零。

ABSTRACT

Motivated by the recent developments in artificial intelligence, we introduce linear quadratic deep structured teams in this paper. Two notions of equivariant and partially equivariant systems are defined, and it is shown that such systems can be partitioned into a few sub-populations of decision makers, where every decision maker in each sub-population is coupled in both dynamics and cost function through a set of linear regressions of the states and actions of all decision makers. Two non-classical information structures are considered: deep-state sharing and partial deep-state sharing, where deep state refers to the linear regression of the states of the decision makers in each sub-population. For a risk-sensitive cost function with deep-state sharing structure, a closed-form low-complexity representation of the globally optimal strategy is obtained, whose computational complexity is independent of the number of decision makers in each sub-population. In addition, it is shown that the risk-sensitive solution converges to the risk-neutral one as the number of decision makers increases to infinity. Moreover, two sub-optimal sequential strategies under partial deep-state sharing information structure are proposed by introducing two Kalman-like filters, one based on the finite-population model and the other one based on the infinite-population model. It is proved that the prices of information associated with the above sub-optimal solutions converge to zero as the number of decision makers goes to infinity. Furthermore, a class of feed-forward deep neural networks with multiple layers of weighted sums and products is introduced wherein the optimal weights and biases are explicitly obtained. A supply-chain management example is presented to demonstrate the efficacy of the obtained results.

研究动机与目标

  • 解决在通信受限和隐私保护等实际约束下,具有大量相互关联决策者的大型网络化系统中的去中心化控制挑战。
  • 为具有连续状态和动作的团队中的线性二次控制问题,开发一种受深度神经网络架构启发的可扩展框架。
  • 制定并求解两种非经典信息结构下的最优控制策略:深度状态共享和部分深度状态共享。
  • 展示所提解决方案在决策者数量增加时的可扩展性和收敛性特性。
  • 建立深度结构化团队与深度前馈神经网络之间的理论联系,明确计算最优权重和偏置。

提出的方法

  • 定义等变和部分等变系统,其中决策者被划分为子群体,通过状态和动作的线性回归相互耦合。
  • 引入两种信息结构:深度状态共享(可完全访问状态的线性回归)和部分深度状态共享(有限访问此类回归)。
  • 应用规范变换和定制的试探解,将哈密顿-雅可比-贝尔曼方程简化为低维的里卡蒂方程组,包括局部和全局里卡蒂方程。
  • 在深度状态共享结构下,推导出计算复杂度与子群体规模无关的闭式全局最优策略。
  • 提出两种次优的顺序策略,采用类似卡尔曼滤波器的方法——一种基于有限群体近似,另一种基于无限群体近似。
  • 证明当决策者数量趋于无穷大时,信息价格和鲁棒性价格均收敛至零。

实验结果

研究问题

  • RQ1在动态和成本函数耦合、信息受限的大型决策者团队中,如何实现最优控制?
  • RQ2当决策者共享深度状态(状态的加权平均)和动作时,最优解的结构是什么?
  • RQ3所提解决方案的规模如何随决策者数量变化?其复杂度能否实现与团队规模无关?
  • RQ4在部分深度状态共享下,次优策略的性能保证如何,特别是在大规模团队的极限情况下?
  • RQ5该框架在结构和参数优化方面,与深度神经网络架构在多大程度上可建立联系?

主要发现

  • 在深度状态共享结构下,推导出闭式全局最优策略,其计算复杂度与每个子群体中的决策者数量无关。
  • 随着决策者数量趋于无穷大,风险敏感最优解收敛至风险中性解。
  • 基于有限群体和无限群体近似的两种次优策略的信息价格,均在决策者数量趋于无穷大时收敛至零。
  • 所提框架生成一类前馈深度神经网络,其最优权重和偏置可通过求解低维里卡蒂方程显式计算得出。
  • 供应链管理案例表明,即使在影响因素不对称的情况下,最优策略仍能有效协调供应商生产与分销商配送。
  • 理论框架在结构和计算层面建立了深度结构化团队与深度前馈神经网络之间的直接类比,尤其体现在层间加权和与乘积的使用上。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。