Skip to main content
QUICK REVIEW

[Paper Review] Camera Based mmWave Beam Prediction: Towards Multi-Candidate Real-World Scenarios

Gouranga Charan, Muhammad Alrabeiah|arXiv (Cornell University)|Aug 14, 2023
Millimeter-Wave Propagation and Modeling4 citations
TL;DR

This paper proposes a vision-aided beam prediction framework using camera and positional data to reduce mmWave beam training overhead in real-world multi-candidate vehicle-to-infrastructure scenarios. Leveraging deep learning, it achieves ~95% top-5 beam prediction accuracy and >93% transmitter identification accuracy on the real-world DeepSense 6G dataset, significantly outperforming prior synthetic-data-based approaches.

ABSTRACT

Leveraging sensory information to aid the millimeter-wave (mmWave) and sub-terahertz (sub-THz) beam selection process is attracting increasing interest. This sensory data, captured for example by cameras at the basestations, has the potential of significantly reducing the beam sweeping overhead and enabling highly-mobile applications. The solutions developed so far, however, have mainly considered single-candidate scenarios, i.e., scenarios with a single candidate user in the visual scene, and were evaluated using synthetic datasets. To address these limitations, this paper extensively investigates the sensing-aided beam prediction problem in a real-world multi-object vehicle-to-infrastructure (V2I) scenario and presents a comprehensive machine learning-based framework. In particular, this paper proposes to utilize visual and positional data to predict the optimal beam indices as an alternative to the conventional beam sweeping approaches. For this, a novel user (transmitter) identification solution has been developed, a key step in realizing sensing-aided multi-candidate and multi-user beam prediction solutions. The proposed solutions are evaluated on the large-scale real-world DeepSense $6$G dataset. Experimental results in realistic V2I communication scenarios indicate that the proposed solutions achieve close to $100\%$ top-5 beam prediction accuracy for the scenarios with single-user and close to $95\%$ top-5 beam prediction accuracy for multi-candidate scenarios. Furthermore, the proposed approach can identify the probable transmitting candidate with more than $93\%$ accuracy across the different scenarios. This highlights a promising approach for nearly eliminating the beam training overhead in mmWave/THz communication systems.

Motivation & Objective

  • Address the gap in real-world evaluation of vision-aided mmWave beam prediction, particularly in dynamic, visually diverse environments with multiple potential transmitters.
  • Overcome the limitations of prior work that relied on synthetic datasets and single-candidate assumptions.
  • Develop a robust, multi-candidate beam prediction system that integrates visual and positional data for accurate beam index selection.
  • Enable practical deployment of sensing-aided beam training in high-mobility mmWave/THz networks by validating on a large-scale real-world dataset.

Proposed method

  • Utilizes RGB camera data and GPS/positioning information from the DeepSense 6G dataset to train a deep neural network (DNN) for beam prediction.
  • Introduces a novel user (transmitter) identification module to resolve the multi-candidate dilemma by distinguishing the actual transmitting vehicle from visual distractors.
  • Employs a multi-task DNN architecture that jointly predicts beam indices and identifies the correct transmitting candidate.
  • Applies data augmentation and domain generalization techniques to improve robustness across diverse real-world scenarios.
  • Uses top-1 and top-5 beam prediction accuracy as evaluation metrics, with power scatter plots to assess practical link performance.
  • Validates the framework on a large-scale real-world dataset collected in urban V2I environments, ensuring practical relevance.

Experimental results

Research questions

  • RQ1Can the performance of vision-aided beam prediction frameworks trained on synthetic data be generalized to real-world, multi-candidate scenarios?
  • RQ2How does the presence of multiple visual candidates affect beam prediction accuracy and system reliability in real-world mmWave communications?
  • RQ3Can a deep learning model effectively identify the correct transmitting vehicle among multiple candidates in complex, dynamic visual scenes?
  • RQ4What is the required amount of training data to achieve robust beam prediction performance in multi-candidate environments?
  • RQ5How does reliance on top-1 beam predictions impact received signal power compared to ground truth beams in multi-candidate settings?

Key findings

  • The proposed framework achieves nearly 100% top-5 beam prediction accuracy in single-user scenarios on the DeepSense 6G dataset.
  • In multi-candidate scenarios, the system maintains 95% top-5 beam prediction accuracy, demonstrating strong generalization to complex visual environments.
  • Transmitter identification accuracy exceeds 93% across all evaluated scenarios, significantly reducing the risk of catastrophic beam misprediction.
  • The confusion matrix shows that most predictions are either correct or near-optimal, indicating robustness even under misprediction.
  • When relying solely on top-1 predictions, the system suffers from a notable performance drop, with R² score decreasing due to low-power beam predictions in ~10% of cases.
  • The power scatter plot reveals that in some cases, the top-1 predicted beam achieves only 10–30% of the received power of the ground truth beam, highlighting the need for lightweight beam training as a fallback.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.