[Paper Review] Cost-Effective Task Offloading Scheduling for Hybrid Mobile Edge-Quantum Computing
This paper proposes a deep reinforcement learning (DRL)-based Lyapunov optimization framework for cost-effective task offloading in hybrid mobile edge-quantum computing (MEQC) systems. By reformulating time-coupled constraints into a penalty-based objective and employing DQN for mode selection and DDPG for partial-task offloading, the approach achieves lower time-average cost and improved sustainability compared to baselines, with optimal performance at Δ=0.7.
In this paper, we aim to address the challenge of hybrid mobile edge-quantum computing (MEQC) for sustainable task offloading scheduling in mobile networks. We develop cost-effective designs for both task offloading mode selection and resource allocation, subject to the individual link latency constraint guarantees for mobile devices, while satisfying the required success ratio for their computation tasks. Specifically, this is a time-coupled offloading scheduling optimization problem in need of a computationally affordable and effective solution. To this end, we propose a deep reinforcement learning (DRL)-based Lyapunov approach. More precisely, we reformulate the original time-coupled challenge into a mixed-integer optimization problem by introducing a penalty part in terms of virtual queues constructed by time-coupled constraints to the objective function. Subsequently, a Deep Q-Network (DQN) is adopted for task offloading mode selection. In addition, we design the Deep Deterministic Policy Gradient (DDPG)-based algorithm for partial-task offloading decision-making. Finally, tested in a realistic network setting, extensive experiment results demonstrate that our proposed approach is significantly more cost-effective and sustainable compared to existing methods.
Motivation & Objective
- To address the challenge of sustainable and cost-effective task offloading scheduling in hybrid mobile edge-quantum computing (MEQC) systems.
- To jointly optimize task offloading mode selection and resource allocation under individual link latency and success ratio constraints.
- To develop a computationally affordable solution for time-coupled offloading scheduling in dynamic MEQC environments.
- To integrate deep reinforcement learning with Lyapunov optimization for real-time, adaptive decision-making.
- To evaluate the proposed method in a realistic network setting and demonstrate superior cost-effectiveness and sustainability.
Proposed method
- Reformulate the time-coupled offloading problem into a mixed-integer optimization problem by introducing a penalty term based on virtual queues constructed from time-coupled constraints.
- Use a Deep Q-Network (DQN) to make discrete decisions on task offloading modes (e.g., local, edge, quantum).
- Employ a Deep Deterministic Policy Gradient (DDPG) algorithm for continuous partial-task offloading decisions.
- Integrate Lyapunov optimization to stabilize queue backlogs and ensure long-term constraint satisfaction.
- Construct a time-average cost objective that balances energy, latency, and resource usage while maintaining QoS guarantees.
- Train the DRL agents in a realistic network environment to learn optimal scheduling policies under dynamic conditions.
Experimental results
Research questions
- RQ1What is the optimal balance between local, edge, and quantum offloading that minimizes long-term cost while satisfying latency and success ratio constraints?
- RQ2How does the proposed DRL-based Lyapunov approach compare to baseline methods in terms of cost-effectiveness and sustainability?
- RQ3What is the impact of the control parameter Δ on the time-average cost and system performance?
- RQ4Can the proposed method achieve stable queue backlogs and constraint satisfaction in time-coupled offloading scenarios?
- RQ5How does the system perform under varying task arrival rates, data sizes, and qubit counts?
Key findings
- The proposed DRL-based Lyapunov algorithm achieves significantly lower time-average cost than baseline methods in realistic network simulations.
- An optimal value of Δ=0.7 minimizes the time-average cost, with performance degrading for values both above and below this point.
- For other baseline methods, the time-average cost remains relatively unchanged across different Δ values, indicating less sensitivity to parameter tuning.
- The system maintains stable queue backlogs and satisfies all latency and success ratio constraints over time.
- The DQN and DDPG components effectively learn to balance offloading decisions, leading to improved energy and resource efficiency.
- Theoretical analysis confirms that the algorithm achieves near-optimal cost performance with bounded queue growth and constraint violations.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.