[Paper Review] Generative AI-aided Joint Training-free Secure Semantic Communications via Multi-modal Prompts
This paper proposes a training-free, generative AI-aided semantic communication system using multi-modal prompts for accurate image reconstruction without joint encoder-decoder training. By leveraging diffusion models and a friendly jammer for covert communication, the framework optimizes power, jamming, and diffusion steps to achieve high SSIM (up to 0.92) and low BEP (<10⁻⁵) under secure transmission conditions.
Semantic communication (SemCom) holds promise for reducing network resource consumption while achieving the communications goal. However, the computational overheads in jointly training semantic encoders and decoders-and the subsequent deployment in network devices-are overlooked. Recent advances in Generative artificial intelligence (GAI) offer a potential solution. The robust learning abilities of GAI models indicate that semantic decoders can reconstruct source messages using a limited amount of semantic information, e.g., prompts, without joint training with the semantic encoder. A notable challenge, however, is the instability introduced by GAI's diverse generation ability. This instability, evident in outputs like text-generated images, limits the direct application of GAI in scenarios demanding accurate message recovery, such as face image transmission. To solve the above problems, this paper proposes a GAI-aided SemCom system with multi-model prompts for accurate content decoding. Moreover, in response to security concerns, we introduce the application of covert communications aided by a friendly jammer. The system jointly optimizes the diffusion step, jamming, and transmitting power with the aid of the generative diffusion models, enabling successful and secure transmission of the source messages.
Motivation & Objective
- To address the high computational and energy cost of joint training in semantic communication systems.
- To enable accurate message reconstruction in semantic communication without requiring joint training of encoders and decoders.
- To enhance security in prompt transmission by integrating covert communication techniques with friendly jamming.
- To design multi-modal prompts combining textual and visual components that preserve structural fidelity while minimizing information leakage.
- To optimize resource allocation (transmit power, jamming, diffusion steps) for reliable and covert transmission using generative diffusion models.
Proposed method
- The system uses multi-modal prompts: textual prompts for semantic content and visual prompts for structural fidelity to guide GAI-based image generation.
- A generative diffusion model (GDM) is employed to reconstruct images from prompts, with denoising applied progressively from Gaussian noise using condition vector c.
- The resource allocation scheme is generated by a policy network ηθ, which predicts optimal transmit power, jamming power, and diffusion steps based on channel conditions.
- A critic network Qυ evaluates the quality of generated images using SSIM as the reward signal, guiding policy optimization via deep reinforcement learning.
- The training uses a loss function minimizing the difference between predicted and actual SSIM values, with Qυ network trained to predict true Q-values.
- Covert communication is achieved by jointly optimizing jamming power and transmission power to maintain low detection probability at the warden while ensuring reliable decoding at the receiver.

Experimental results
Research questions
- RQ1Can generative AI enable accurate semantic communication without joint training of encoders and decoders?
- RQ2How can multi-modal prompts improve image reconstruction fidelity in semantic communication systems?
- RQ3What is the impact of prompt instability in generative models on semantic communication reliability?
- RQ4How can covert communication be effectively integrated into GAI-aided semantic communication to prevent eavesdropping?
- RQ5What resource allocation strategy maximizes image reconstruction quality while maintaining low detection probability in the presence of a warden?
Key findings
- The GDM-based resource allocation scheme outperforms DRL-based methods in test rewards, indicating better optimization of transmission parameters.
- With a BEP of 10⁻⁵, approximately 50 diffusion steps are sufficient to achieve high-quality image reconstruction with SSIM above 0.92.
- As BEP increases to 10⁻³, reconstructed images become noisy and SSIM scores degrade significantly, highlighting the need for low-error transmission.
- Jamming power above 24 dBW ensures detection probability exceeds the threshold, satisfying covert communication requirements under the given environment.
- The system achieves a covert rate that decreases with increasing jamming power, balancing security and data rate.
- The proposed method enables high-fidelity image reconstruction without joint training, reducing computational and energy costs in semantic communication.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.