[Paper Review] Learning Aided Optimization for Energy Harvesting Devices with Outdated State Information
This paper proposes a learning-aided power control algorithm for energy harvesting devices that operates with outdated system state information and unknown probability distributions. By combining online gradient learning with drift-plus-penalty optimization, the algorithm achieves utility within $O(\epsilon)$ of optimality using a battery of size $O(1/\epsilon)$, with convergence time $O(1/\epsilon^2)$, enabling low-complexity, adaptive control without requiring real-time state knowledge or distributional assumptions.
This paper considers utility optimal power control for energy harvesting wireless devices with a finite capacity battery. The distribution information of the underlying wireless environment and harvestable energy is unknown and only outdated system state information is known at the device controller. This scenario shares similarity with Lyapunov opportunistic optimization and online learning but is different from both. By a novel combination of Zinkevich's online gradient learning technique and the drift-plus-penalty technique from Lyapunov opportunistic optimization, this paper proposes a learning-aided algorithm that achieves utility within $O(ε)$ of the optimal, for any desired $ε>0$, by using a battery with an $O(1/ε)$ capacity. The proposed algorithm has low complexity and makes power investment decisions based on system history, without requiring knowledge of the system state or its probability distribution.
Motivation & Objective
- To address the challenge of utility-optimal power control in energy harvesting devices when neither the current system state nor its distribution is known.
- To design a low-complexity control policy that operates using only outdated system state information and historical data.
- To close the gap in existing methods that require either full state knowledge or i.i.d. assumptions, especially in dynamic, non-stationary environments.
- To establish a theoretical performance tradeoff between utility optimality and battery size, showing that $O(\epsilon)$ optimality is achievable with $O(1/\epsilon)$ battery capacity.
Proposed method
- The algorithm integrates Zinkevich’s online gradient learning with the drift-plus-penalty (DPP) framework from Lyapunov optimization to handle time-varying constraints and unknown distributions.
- It uses a time-varying constraint set $\mathcal{P}(t)$ that limits power usage to the current battery level $E[t]$, ensuring energy causality.
- The power update rule incorporates a dual variable $Q[t]$ and a penalty parameter $V$, balancing utility maximization and battery stability.
- A projection step ensures that the power vector $\mathbf{p}[t]$ remains within feasible bounds at each slot, even when energy is insufficient.
- The algorithm handles delayed state observations by modifying the gradient update to use past system states, preserving performance guarantees.
- The method is extended to non-i.i.d. system states by modeling state evolution via a Markov chain, maintaining the same performance tradeoffs.
Experimental results
Research questions
- RQ1Can a learning-aided algorithm achieve near-optimal utility in energy harvesting systems when only outdated system state information is available and the underlying distribution is unknown?
- RQ2What is the fundamental tradeoff between utility optimality and battery capacity in such systems?
- RQ3How does observation delay affect the convergence and performance of online learning-based power control in energy harvesting devices?
- RQ4Can the proposed algorithm outperform baseline methods that use either simple projection or outdated utility maximization under energy constraints?
- RQ5Does the performance guarantee hold under non-i.i.d. system state processes, such as Markov-modulated channels?
Key findings
- The proposed algorithm achieves utility within $O(\epsilon)$ of the optimal utility for any $\epsilon > 0$, demonstrating near-optimality under unknown system dynamics.
- A battery capacity of $O(1/\epsilon)$ is sufficient to achieve $O(\epsilon)$ optimality, establishing a tight tradeoff between performance and hardware cost.
- The convergence time of the algorithm is $O(1/\epsilon^2)$, which is polynomial in the desired accuracy $\epsilon$, ensuring practical convergence.
- Simulation results show that the algorithm outperforms two baselines: one using projected online gradient learning and another using outdated utility maximization.
- Even with observation delays of up to 10 slots, the algorithm maintains long-term performance with only a minor impact on convergence speed, confirming robustness to delayed feedback.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.