Skip to main content
QUICK REVIEW

[Paper Review] Constrained Contextual Bandit Learning for Adaptive Radar Waveform Selection

Charles E. Thornton, R. Michael Buehrer|arXiv (Cornell University)|Mar 9, 2021
Radar Systems and Signal ProcessingEngineering55 references31 citations
TL;DR

This paper formulates adaptive radar waveform selection as a constrained linear contextual bandit problem, enabling sample-efficient, online learning in dynamic environments. By integrating spectrum sensing and receiver feedback, it uses Thompson Sampling and EXP3 algorithms to improve target detection and reduce Doppler sidelobes via a time-varying waveform distortion constraint, outperforming fixed and naive adaptive radar in coexistence and jamming scenarios.

ABSTRACT

A sequential decision process in which an adaptive radar system repeatedly interacts with a finite-state target channel is studied. The radar is capable of passively sensing the spectrum at regular intervals, which provides side information for the waveform selection process. The radar transmitter uses the sequence of spectrum observations as well as feedback from a collocated receiver to select waveforms which accurately estimate target parameters. It is shown that the waveform selection problem can be effectively addressed using a linear contextual bandit formulation in a manner that is both computationally feasible and sample efficient. Stochastic and adversarial linear contextual bandit models are introduced, allowing the radar to achieve effective performance in broad classes of physical environments. Simulations in a radar-communication coexistence scenario, as well as in an adversarial radar-jammer scenario, demonstrate that the proposed formulation provides a substantial improvement in target detection performance when Thompson Sampling and EXP3 algorithms are used to drive the waveform selection process. Further, it is shown that the harmful impacts of pulse-agile behavior on coherently processed radar data can be mitigated by adopting a time-varying constraint on the radar's waveform catalog.

Motivation & Objective

  • To address the challenge of real-time, adaptive radar waveform selection in interference-limited and dynamic environments with limited prior knowledge.
  • To develop a computationally feasible and sample-efficient online learning framework for radar waveform adaptation.
  • To mitigate harmful Doppler sidelobe effects from pulse-agile waveform selection using a time-varying distortion constraint.
  • To enable robust performance in both stochastic and adversarial environments, including radar-jamming and radar-cellular coexistence scenarios.
  • To integrate online learning with radar tracking systems for improved overall sensing performance.

Proposed method

  • Formulates radar waveform selection as a linear contextual bandit problem using spectrum observations as context and receiver feedback as reward.
  • Introduces a time-varying constraint on waveform transitions based on a distortion metric between consecutive waveforms to reduce Doppler sidelobes.
  • Employs Thompson Sampling and EXP3 algorithms for online decision-making under uncertainty in both stochastic and adversarial environments.
  • Uses a linear reward model where the expected reward depends on a context vector combining spectrum sensing and target channel state.
  • Applies regret bounds from online learning theory to guarantee performance under both Bayesian and frequentist assumptions.
  • Integrates the learned waveform policy with a Kalman tracker to improve range and Doppler estimation accuracy.

Experimental results

Research questions

  • RQ1Can a contextual bandit framework effectively model adaptive radar waveform selection under limited feedback and dynamic interference?
  • RQ2How does incorporating a time-varying waveform distortion constraint impact Doppler sidelobe levels and detection performance?
  • RQ3What performance gains does online learning with Thompson Sampling and EXP3 provide over fixed or naive adaptive radar in coexistence and jamming scenarios?
  • RQ4How do the theoretical regret bounds of the proposed algorithms scale with context dimension and time horizon?
  • RQ5To what extent can the constrained contextual bandit model improve tracking accuracy in real-world radar deployments?

Key findings

  • The constrained EXP3 algorithm reduced average cost below 0.1 in the unconstrained case and improved detection performance by making the target peak more pronounced and sidelobes lower.
  • With a distortion constraint of ˆd=0.2, Doppler sidelobe levels were significantly reduced, and the true target peak became clearly visible in range-Doppler images.
  • The constrained EXP3 algorithm achieved lower tracking RMSE than unconstrained variants, indicating improved range and Doppler estimation accuracy.
  • Thompson Sampling with the distortion constraint also improved detection performance, demonstrating robustness in adverse conditions.
  • The proposed framework achieved favorable performance in both radar-cellular coexistence and intentional jamming scenarios, outperforming fixed-bandwidth radar and naive adaptive schemes.
  • Theoretical regret bounds for Thompson Sampling and EXP3 were derived, showing O(√n log n) and O(√n) scaling under appropriate assumptions, respectively.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.