[Paper Review] Deep Reinforcement Learning for Intelligent Reflecting Surface-assisted D2D Communications
This paper proposes a deep reinforcement learning (DRL) framework using proximal policy optimization (PPO) to jointly optimize transmit power at D2D transmitters and phase shifts at an intelligent reflecting surface (IRS), significantly improving network sum-rate in IRS-assisted D2D communications. The method achieves superior performance over benchmark schemes, especially under high interference and QoS constraints, with low computational overhead.
In this paper, we propose a deep reinforcement learning (DRL) approach for solving the optimisation problem of the network's sum-rate in device-to-device (D2D) communications supported by an intelligent reflecting surface (IRS). The IRS is deployed to mitigate the interference and enhance the signal between the D2D transmitter and the associated D2D receiver. Our objective is to jointly optimise the transmit power at the D2D transmitter and the phase shift matrix at the IRS to maximise the network sum-rate. We formulate a Markov decision process and then propose the proximal policy optimisation for solving the maximisation game. Simulation results show impressive performance in terms of the achievable rate and processing time.
Motivation & Objective
- To address the challenge of interference and low spectral efficiency in dense D2D communications by leveraging intelligent reflecting surfaces (IRS).
- To jointly optimize transmit power at D2D transmitters and phase shift matrices at the IRS to maximize network sum-rate.
- To overcome the limitations of conventional optimization methods, such as high computational complexity and reliance on perfect CSI, by using deep reinforcement learning.
- To design a scalable, real-time resource allocation solution that adapts to dynamic network conditions without requiring full channel state information.
Proposed method
- Formulates the joint power and phase shift optimization as a Markov decision process (MDP), modeling the D2D-IRS system as a sequential decision-making problem.
- Employs the proximal policy optimization (PPO) algorithm with a clipping surrogate objective to stabilize training and improve sample efficiency.
- Uses a deep neural network to represent the policy network, mapping state observations (e.g., user locations, CSI estimates) to actions (transmit power and phase shifts).
- Defines the environment reward as the network sum-rate, encouraging the agent to maximize spectral efficiency while satisfying QoS constraints.
- Integrates a clipping mechanism in the PPO algorithm to prevent large policy updates and ensure stable learning.
- Trains the DRL agent offline using simulated interactions with the environment, enabling real-time deployment for online control.
Experimental results
Research questions
- RQ1Can deep reinforcement learning effectively solve the joint power and phase shift optimization problem in IRS-assisted D2D networks?
- RQ2How does the proposed PPO-based DRL method compare to conventional optimization and heuristic schemes in terms of sum-rate and convergence speed?
- RQ3To what extent does the DRL approach maintain performance under imperfect or partial channel state information?
- RQ4How does the system performance scale with increasing numbers of D2D pairs and IRS elements?
- RQ5What is the impact of QoS constraints on the achievable sum-rate when using DRL-based optimization?
Key findings
- The proposed PPO-based DRL scheme achieves the highest network sum-rate across all tested configurations, outperforming MPT, RPS, and no-IRS baselines.
- With 20 IRS elements and 5 D2D pairs, the proposed method achieves a 25–35% higher sum-rate than the MPT baseline and over 50% higher than the RPS and no-IRS schemes.
- As the number of IRS elements increases from 10 to 40, the sum-rate of the proposed method grows monotonically, demonstrating scalability and interference suppression capability.
- When the number of D2D pairs exceeds 6, the proposed method maintains increasing sum-rate, while other schemes degrade due to interference overload.
- Under high QoS thresholds (e.g., r_min ≥ 7), the proposed method maintains a significant performance gap over MPT, which degrades sharply, highlighting its robustness.
- At high transmit power (400 mW), the performance gap between the proposed method and baselines widens, confirming the effectiveness of joint optimization in interference-limited scenarios.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.