Skip to main content
QUICK REVIEW

[论文解读] Learning-based attacks in Cyber-Physical Systems: Exploration, Detection, and Control Cost trade-offs

Anshuka Rangi, Mohammad Javad Khojasteh|arXiv (Cornell University)|Nov 21, 2020
Smart Grid Security and Resilience参考文献 50被引用 6
一句话总结

本文研究线性网络物理系统中的基于学习的中间人(MITM)攻击,其中攻击者在探索阶段学习系统动态,并在后续阶段利用这些信息欺骗控制器。研究建立了攻击检测时间、学习时长与控制能量之间的紧密权衡关系,证明了在置信度约束下,可靠检测所需的欺骗时间与控制能耗的阶最优下界。

ABSTRACT

We study the problem of learning-based attacks in linear systems, where the communication channel between the controller and the plant can be hijacked by a malicious attacker. We assume the attacker learns the dynamics of the system from observations, then overrides the controller's actuation signal, while mimicking legitimate operation by providing fictitious sensor readings to the controller. On the other hand, the controller is on a lookout to detect the presence of the attacker and tries to enhance the detection performance by carefully crafting its control signals. We study the trade-offs between the information acquired by the attacker from observations, the detection capabilities of the controller, and the control cost. Specifically, we provide tight upper and lower bounds on the expected $ε$-deception time, namely the time required by the controller to make a decision regarding the presence of an attacker with confidence at least $(1-ε\log(1/ε))$. We then show a probabilistic lower bound on the time that must be spent by the attacker learning the system, in order for the controller to have a given expected $ε$-deception time. We show that this bound is also order optimal, in the sense that if the attacker satisfies it, then there exists a learning algorithm with the given order expected deception time. Finally, we show a lower bound on the expected energy expenditure required to guarantee detection with confidence at least $1-ε\log(1/ε)$.

研究动机与目标

  • 理解基于学习的中间人攻击在何种条件下可在线性网络物理系统中保持隐蔽。
  • 分析攻击者学习阶段持续时间、控制器检测置信度与控制能量成本之间的权衡关系。
  • 推导出可靠检测所需欺骗时间与控制能耗的理论下界。
  • 证明这些下界为阶最优,即可通过特定的学习与检测策略实现。
  • 量化攻击者为实现给定期望欺骗时间并以高置信度达成目标,所需最小学习时间。

提出的方法

  • 将攻击建模为两阶段过程:探索阶段(通过窃听学习系统动态)与利用阶段(使用虚假反馈欺骗控制器)。
  • 采用统计决策框架,控制器基于观测反馈信号的差异性检测攻击者。
  • 利用信息论工具,结合形如 1−ϵlog(1/ϵ) 的置信度约束,推导出期望 ϵ-欺骗时间的下界。
  • 建立攻击者探索阶段持续时间的随机下界,表明其必须随 D/log(1/ϵ) 的阶增长,其中 D 为给定的欺骗时间。
  • 通过构造一种学习算法,证明阶最优性:当探索时间达到 O(D/log(1/ϵ)) 时,可实现至少 D 的期望欺骗时间。
  • 推导出为在时间 D 内以置信度 1−ϵlog(1/ϵ) 保证检测攻击者,所需控制能量的下界,表明其随 Ω(D/log(1/ϵ)) 增长。

实验结果

研究问题

  • RQ1控制器以置信度 1−ϵlog(1/ϵ) 检测基于学习的中间人攻击者,所需最小期望时间是多少?
  • RQ2攻击者为以高置信度实现目标欺骗时间 D,必须在探索阶段花费多长学习时间?
  • RQ3是否存在确保在给定时间范围内检测到攻击者的控制能量支出的理论下界?
  • RQ4所推导出的欺骗时间与控制成本下界在实际中是否可实现?是否为阶最优?
  • RQ5攻击者的学习精度、控制器的检测能力与系统控制成本之间存在何种关系?

主要发现

  • 期望 ϵ-欺骗时间受到攻击者学习算法参数与控制器检测策略的下界约束,确立了可检测性的根本极限。
  • 攻击者探索阶段持续时间的随机下界为 Ω(D/log(1/ϵ)),当 ϵ→0 时,表明更长的欺骗时间需要更长的学习时间。
  • 该下界为阶最优:若攻击者学习时间达到 O(D/log(1/ϵ)),则存在一种学习算法可实现至少 D 的期望欺骗时间。
  • 为在时间 D 内以置信度 1−ϵlog(1/ϵ) 保证检测攻击者,所需期望控制能量至少为 Ω(D/log(1/ϵ)),表明控制成本与检测速度之间存在直接权衡。
  • 研究结果建立了学习时间、欺骗时间与控制能量之间的紧密、阶最优的权衡关系,为安全网络物理系统设计提供了理论基础。
  • 该框架适用于广泛的学习算法与检测策略,展示了在基于学习的攻击背景下具有鲁棒性与普适性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。