Skip to main content
QUICK REVIEW

[论文解读] Wide Area Measurement System-based Low Frequency Oscillation Damping Control through Reinforcement Learning

Yousaf Hashmy, Zhe Yu|arXiv (Cornell University)|Jan 22, 2020
Power System Optimization and Stability参考文献 25被引用 5
一句话总结

本文提出了一种基于强化学习(RL)的广域阻尼控制策略,利用相量测量单元(PMU)数据,在通信延迟和可再生能源不确定性条件下,有效缓解电力系统中的低频振荡。通过采用具有策略梯度优化的深度确定性策略梯度(DDPG)方法,该方法在多种通信信道和高比例可再生能源接入场景下,实现了鲁棒、可扩展且可解释的阻尼控制,其性能优于传统PID和网络化预测控制方法。

ABSTRACT

Ensuring the stability of power systems is gaining more attraction today than ever before, due to the rapid growth of uncertainties in load and renewable energy penetration. Lately, wide area measurement system-based centralized controlling techniques started providing a more flexible and robust control to keep the system stable. But, such a modernization of control philosophy faces pressing challenges due to the irregularities in delays of long-distance communication channels and response of equipment to control actions. Therefore, we propose an innovative approach that can revolutionize the control strategy for damping down low frequency oscillations in transmission systems. Proposed method is enriched with a potential of overcoming the challenges of communication delays and other non-linearities in wide area damping control by leveraging the capability of the reinforcement learning technique. Such a technique has a unique characteristic to learn on diverse scenarios and operating conditions by exploring the environment and devising an optimal control action policy by implementing policy gradient method. Our detailed analysis and systematically designed numerical validation prove the feasibility, scalability and interpretability of the carefully modelled low-frequency oscillation damping controller so that stability is ensured even with the uncertainties of load and generation are on the rise.

研究动机与目标

  • 应对由于可再生能源大规模接入和负荷不确定性增加,导致现代电力系统中低频振荡问题日益严峻的挑战。
  • 克服传统集中式广域阻尼控制的局限性,特别是对可变通信延迟和设备响应非线性特性的敏感性。
  • 开发一种鲁棒、自适应的控制策略,通过与动态系统环境的交互学习最优控制策略,而无需依赖精确的系统模型。
  • 确保在多种通信信道(微波、光纤、卫星、电力线载波)和随机延迟条件下,系统的稳定性和可观测性。
  • 验证基于强化学习的控制策略在大规模非线性电力系统中高比例可再生能源渗透场景下的可扩展性和实时可行性。

提出的方法

  • 利用广域测量系统(WAMS)中PMU提供的数据,为控制决策提供实时、全系统范围的状态信息。
  • 采用深度确定性策略梯度(DDPG)强化学习方法,学习一种将系统状态映射到最优控制动作的随机控制策略。
  • 应用策略梯度优化,并使用深度神经网络作为函数逼近器,以处理高维状态空间中的连续控制动作。
  • 将通信延迟建模为随机过程,采用高斯混合模型模拟远距离信号传输中的真实不确定性。
  • 设计一种奖励函数,对振荡幅度和偏离额定频率的程度进行惩罚,以鼓励快速阻尼和系统稳定性。
  • 在不同可再生能源渗透水平和通信信道类型下,使用多机电力系统模型对控制器进行验证。

实验结果

研究问题

  • RQ1在可变通信延迟条件下,基于强化学习的控制器是否能有效阻尼大规模电力系统中的低频振荡?
  • RQ2在高比例可再生能源渗透和不确定延迟条件下,所提出的RL控制器与传统PID和网络化预测控制(NPC)相比表现如何?
  • RQ3RL智能体在探索多样化运行条件和随机延迟的过程中,能在多大程度上学习到鲁棒的控制策略?
  • RQ4基于DDPG的控制器是否能在包括微波、光纤、卫星和电力线载波在内的不同通信介质上保持稳定性和性能?
  • RQ5当控制器在不确定延迟与恒定延迟条件下学习时,训练时间与鲁棒性之间的权衡如何?

主要发现

  • 基于DDPG的强化学习控制器在高比例可再生能源渗透条件下,成功阻尼了所有通信信道(微波、光纤、卫星和电力线载波)中的低频振荡。
  • 与PID和NPC控制器不同,后者在电力线载波延迟下因振荡持续增长而无法稳定系统,该RL控制器保持了系统稳定并实现了完全阻尼。
  • 在将可变时间延迟建模为高斯混合分布的场景下,RL智能体虽需更多训练回合,但相比固定延迟场景,表现出更优的鲁棒性和泛化能力。
  • 在所有测试场景中,RL控制器均优于PID和NPC控制器,尤其在高不确定性条件下表现更优:在微波和光纤链路中阻尼速度更快,在卫星和电力线载波信道中保持稳定性能。
  • 该方法展现出良好的可扩展性和可解释性,在通信延迟和可再生能源发电不确定性显著时,仍能维持系统可观测性和稳定性。
  • 结果证实,基于强化学习的控制器对非线性特性、模型不确定性及随机延迟具有强鲁棒性,适用于实际广域阻尼控制应用。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。