[Paper Review] Transformer-aided Wireless Image Transmission with Channel Feedback
This paper proposes JSCCformer-f, a Transformer-based joint source-channel coding framework for wireless image transmission that leverages receiver feedback to iteratively refine image reconstruction. By using a unified encoder that fuses semantic image features with feedback-derived belief states, the scheme achieves state-of-the-art performance with low complexity and robustness to feedback noise, outperforming separation-based and prior feedback-aided methods across diverse SNR and bandwidth conditions.
This paper presents a novel wireless image transmission paradigm that can exploit feedback from the receiver, called DeepJSCC-ViT-f. We consider a block feedback channel model, where the transmitter receives noiseless/noisy channel output feedback after each block. The proposed scheme employs a single encoder to facilitate transmission over multiple blocks, refining the receiver's estimation at each block. Specifically, the unified encoder of DeepJSCC-ViT-f can leverage the semantic information from the source image, and acquire channel state information and the decoder's current belief about the source image from the feedback signal to generate coded symbols at each block. Numerical experiments show that our DeepJSCC-ViT-f scheme achieves state-of-the-art transmission performance with robustness to noise in the feedback link. Additionally, DeepJSCC-ViT-f can adapt to the channel condition directly through feedback without the need for separate channel estimation. We further extend the scope of the DeepJSCC-ViT-f approach to include the broadcast channel, which enables the transmitter to generate broadcast codes in accordance with signal semantics and channel feedback from individual receivers.
Motivation & Objective
- Address the high complexity and inadaptability of existing feedback-aided deep JSCC schemes for image transmission.
- Overcome the limitations of separate source and channel coding in finite blocklength regimes by enabling joint optimization with feedback.
- Develop a unified, scalable encoder that leverages semantic information and receiver feedback to iteratively improve image reconstruction.
- Enable channel adaptation without explicit channel estimation by directly using feedback to guide transmission.
- Generalize the framework to broadcast channels, allowing adaptive transmission to multiple receivers with varying channel conditions.
Proposed method
- Employ a single, unified encoder that processes the source image and feedback signals (decoder’s current belief) at each block to generate coded symbols.
- Use a vision Transformer (ViT) with self-attention and cross-attention mechanisms to model long-range dependencies and fuse semantic and feedback information.
- Integrate channel state information and receiver belief from feedback into the encoder’s input via a learnable fusion module.
- Train the end-to-end system using a weighted MSE loss that balances reconstruction quality across multiple receivers in the broadcast setting.
- Leverage block feedback where the transmitter receives noiseless/noisy channel output feedback after each transmission block to refine subsequent transmissions.
- Extend the framework to broadcast channels by training a single encoder to generate signals tailored to multiple receivers based on their individual feedback and channel conditions.
Experimental results
Research questions
- RQ1Can a unified Transformer-based encoder outperform multi-stage, independent encoder-decoder architectures in feedback-aided image transmission?
- RQ2To what extent can feedback improve reconstruction quality without requiring explicit channel estimation?
- RQ3How does the proposed scheme perform under noisy feedback links compared to existing feedback-aided JSCC methods?
- RQ4Can the framework be generalized to multi-user broadcast scenarios while maintaining performance and adaptability?
- RQ5Does the integration of semantic information and receiver belief via attention mechanisms lead to better reconstruction than conventional feature-based approaches?
Key findings
- JSCCformer-f achieves state-of-the-art PSNR performance across all SNR and bandwidth ratio values tested, significantly outperforming DeepJSCC-f and separation-based schemes.
- The scheme improves reconstruction quality by at least 3 dB for each receiver in the broadcast scenario compared to the BPG-Capacity benchmark.
- Average PSNR gains of at least 0.9 dB are observed across all SNR pairs in the broadcast channel evaluation, demonstrating consistent superiority.
- The framework is robust to noisy feedback, maintaining high performance without requiring additional error protection on the feedback link.
- The unified encoder design reduces training and memory complexity compared to training separate encoders for each feedback block.
- The system adapts effectively to varying channel conditions without a separate channel estimation module, relying solely on feedback to guide transmission.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.