[论文解读] Online Regulation of Unstable LTI Systems from a Single Trajectory
本文提出数据引导调节(DGR),一种新颖的在线控制方法,用于仅从单一轨迹调节不稳定线性时不变(LTI)系统,且无需事先了解系统稳定性。通过利用输入矩阵的访问权限和一种新概念‘可正则化性’,DGR 在有限时间内实现系统稳定,同时生成可用于后续系统辨识或稳定化的信息性数据,并与系统的不稳定性数和谱特性建立理论联系。
Recently, data-driven methods for control of dynamic systems have received considerable attention in system theory and machine learning as they provide a mechanism for feedback synthesis from the observed time-series data. However learning, say through direct policy updates, often requires assumptions such as knowing a priori that the initial policy (gain) is stabilizing, e.g., when the open-loop system is stable. In this paper, we examine online regulation of (possibly unstable) partially unknown linear systems with no a priori assumptions on the initial controller. First, we introduce and characterize the notion of ''regularizability'' for linear systems that gauges the capacity of a system to be regulated in finite-time in contrast to its asymptotic behavior (commonly characterized by stabilizability/controllability). Next, having access only to the input matrix, we propose the Data-GuidedRegulation (DGR) synthesis that--as its name suggests--regulates the underlying states while also generating informative data that can subsequently be used for data-driven stabilization or system identification (sysID). The analysis is also related in spirit, to thespectrum and the ''instability number'' of the underlying linear system, a novel geometric property studied in this work. We further elucidate our results by considering special structures for system parameters as well as boosting the performance of the algorithm via a rank-one matrix update using the discrete nature of data collection in the problem setup. Finally, we demonstrate the utility of the proposed approach via an example involving direct (online) regulation of the X-29 aircraft.
研究动机与目标
- 解决在缺乏稳定初始控制器的情况下调节不稳定LTI系统的挑战。
- 定义并表征一种新的‘可正则化性’概念,以捕捉有限时间调节能力,该能力与渐近可稳定性的概念相区别。
- 开发一种数据驱动的在线调节框架,同时实现系统稳定和生成未来可用的有用信息数据。
- 建立所提方法与系统谱特性之间的理论联系,特别是与‘不稳定性数’的关系。
提出的方法
- 提出数据引导调节(DGR)算法,仅通过访问输入矩阵实现实时从单一观测轨迹调节系统状态。
- 引入‘可正则化性’概念作为有限时间调节特性,将传统可稳定性的概念推广至包含瞬态性能。
- 设计DGR合成过程,使其在调节过程中生成信息性数据,从而支持后续的数据驱动系统辨识或稳定化。
- 通过一种新颖的几何特性——‘不稳定性数’——分析该方法,将其与系统的谱特性和调节性能联系起来。
- 采用秩一矩阵更新技术,利用离散数据采集结构以提升算法性能。
- 在真实世界的X-29飞机控制案例中验证该方法,证明其实际应用价值。
实验结果
研究问题
- RQ1在缺乏对系统稳定性或初始稳定控制器的先验知识的情况下,是否能够实现不稳定LTI系统在有限时间内的调节?
- RQ2线性系统的哪些结构和几何特性决定了其有限时间调节能力?
- RQ3如何设计在线调节方法,使其在实现系统稳定的同时,生成对未来系统辨识或控制综合有用的数据?
- RQ4系统的谱特性以及新定义的‘不稳定性数’在实现或限制调节性能方面起什么作用?
- RQ5在离散时间数据采集设置中,是否可通过结构化矩阵更新提升性能?
主要发现
- 所提出的DGR方法成功在有限时间内调节不稳定LTI系统,无需依赖稳定初始控制器,仅依赖对输入矩阵的访问。
- ‘可正则化性’概念为系统有限时间调节能力提供了新的表征方式,与经典可稳定性的概念相区别。
- 该算法在调节过程中生成信息性数据,支持后续基于数据的系统辨识或稳定化,且样本效率更高。
- ‘不稳定性数’——一种系统的新几何特性——被证明是调节性能和收敛速度的关键决定因素。
- 秩一矩阵更新在离散时间数据采集中提升了算法的收敛性和性能,已在X-29飞机案例中得到验证。
- X-29案例研究证明了该方法在实际复杂不稳定系统在线调节中的可行性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。