[论文解读] Constrained Contextual Bandit Learning for Adaptive Radar Waveform Selection
本文将自适应雷达波形选择建模为一个约束性线性上下文Bandit问题,实现了在动态环境中样本高效的在线学习。通过整合频谱感知与接收机反馈,结合Thompson Sampling与EXP3算法,利用时变波形失真约束,提升了目标检测性能并降低了Doppler旁瓣,优于固定波形与简单自适应雷达在共存与干扰场景下的表现。
A sequential decision process in which an adaptive radar system repeatedly interacts with a finite-state target channel is studied. The radar is capable of passively sensing the spectrum at regular intervals, which provides side information for the waveform selection process. The radar transmitter uses the sequence of spectrum observations as well as feedback from a collocated receiver to select waveforms which accurately estimate target parameters. It is shown that the waveform selection problem can be effectively addressed using a linear contextual bandit formulation in a manner that is both computationally feasible and sample efficient. Stochastic and adversarial linear contextual bandit models are introduced, allowing the radar to achieve effective performance in broad classes of physical environments. Simulations in a radar-communication coexistence scenario, as well as in an adversarial radar-jammer scenario, demonstrate that the proposed formulation provides a substantial improvement in target detection performance when Thompson Sampling and EXP3 algorithms are used to drive the waveform selection process. Further, it is shown that the harmful impacts of pulse-agile behavior on coherently processed radar data can be mitigated by adopting a time-varying constraint on the radar's waveform catalog.
研究动机与目标
- 为解决在先验知识有限的干扰受限与动态环境中实时自适应雷达波形选择的挑战。
- 开发一种计算可行且样本高效的雷达波形自适应在线学习框架。
- 通过时变失真约束缓解脉冲敏捷波形选择带来的有害Doppler旁瓣效应。
- 在随机与对抗性环境中实现鲁棒性能,包括雷达-干扰与雷达-蜂窝共存场景。
- 将在线学习与雷达跟踪系统集成,以提升整体感知性能。
提出的方法
- 将雷达波形选择建模为线性上下文Bandit问题,利用频谱观测作为上下文,接收机反馈作为奖励。
- 基于连续波形间失真度量,引入波形跳变的时变约束,以降低Doppler旁瓣。
- 在随机与对抗性环境中,采用Thompson Sampling与EXP3算法进行不确定性下的在线决策。
- 采用线性奖励模型,其中期望奖励取决于结合频谱感知与目标信道状态的上下文向量。
- 应用在线学习理论中的遗憾界,以在贝叶斯与频率学假设下保证性能。
- 将学习到的波形策略与卡尔曼跟踪器集成,以提高距离与Doppler估计精度。
实验结果
研究问题
- RQ1在反馈受限与动态干扰条件下,上下文Bandit框架能否有效建模自适应雷达波形选择?
- RQ2引入时变波形失真约束对Doppler旁瓣电平与检测性能有何影响?
- RQ3在共存与干扰场景下,Thompson Sampling与EXP3的在线学习相比固定波形或简单自适应雷达能带来多大性能提升?
- RQ4所提算法的理论遗憾界在上下文维度与时间跨度上如何变化?
- RQ5约束性上下文Bandit模型在实际雷达部署中能在多大程度上提升跟踪精度?
主要发现
- 约束性EXP3算法在无约束情况下平均代价低于0.1,且通过使目标主瓣更突出、旁瓣更低,提升了检测性能。
- 当失真约束为ˆd=0.2时,Doppler旁瓣电平显著降低,真实目标主瓣在距离-Doppler图像中清晰可见。
- 约束性EXP3算法的跟踪均方根误差(RMSE)低于无约束变体,表明距离与Doppler估计精度得到提升。
- 采用失真约束的Thompson Sampling也改善了检测性能,证明了其在恶劣条件下的鲁棒性。
- 所提框架在雷达-蜂窝共存与有意干扰场景中均表现出优异性能,优于固定带宽雷达与简单自适应方案。
- 推导出Thompson Sampling与EXP3的理论遗憾界,分别在合理假设下呈现O(√n log n)与O(√n)的量级。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。