Skip to main content
QUICK REVIEW

[论文解读] Joint Time Scheduling and Transaction Fee Selection in Blockchain-based RF-Powered Backscatter Cognitive Radio Network

Tran The Anh, Nguyen Cong Luong|arXiv (Cornell University)|Jan 10, 2020
Energy Harvesting in Wireless Networks被引用 6
一句话总结

本文提出了一种基于深度强化学习(DRL)的框架,用于在射频供电背向散射认知无线电网络中联合优化时间调度、区块链网络选择和交易费用。通过建立随机优化问题并采用双值双重深度Q网络(D3QN),网关动态分配时隙用于能量采集、背向散射和传输,同时选择最优区块链和费用率,实现每帧高达9.1个数据单位的吞吐量,显著优于Q-learning和基线方法。

ABSTRACT

In this paper, we develop a new framework called blockchain-based Radio Frequency (RF)-powered backscatter cognitive radio network. In the framework, IoT devices as secondary transmitters transmit their sensing data to a secondary gateway by using the RF-powered backscatter cognitive radio technology. The data collected at the gateway is then sent to a blockchain network for further verification, storage and processing. As such, the framework enables the IoT system to simultaneously optimize the spectrum usage and maximize the energy efficiency. Moreover, the framework ensures that the data collected from the IoT devices is verified, stored and processed in a decentralized but in a trusted manner. To achieve the goal, we formulate a stochastic optimization problem for the gateway under the dynamics of the primary channel, the uncertainty of the IoT devices, and the unpredictability of the blockchain environment. In the problem, the gateway jointly decides (i) the time scheduling, i.e., the energy harvesting time, backscatter time, and transmission time, among the IoT devices, (ii) the blockchain network, and (iii) the transaction fee rate to maximize the network throughput while minimizing the cost. To solve the stochastic optimization problem, we then propose to employ, evaluate, and assess the Deep Reinforcement Learning (DRL) with Dueling Double Deep Q-Networks (D3QN) to derive the optimal policy for the gateway. The simulation results clearly show that the proposed solution outperforms the conventional baseline approaches such as the conventional Q-Learning algorithm and non-learning algorithms in terms of network throughput and convergence speed. Furthermore, the proposed solution guarantees that the data is stored in the blockchain network at a reasonable cost.

研究动机与目标

  • 为解决密集物联网部署中的频谱稀缺和能量约束问题,采用射频供电背向散射认知无线电技术。
  • 通过集成区块链实现去中心化、可信的数据存储和验证,提升数据完整性、透明性和安全性。
  • 在动态、不确定且不可预测的系统条件下,联合优化时间调度(能量采集、背向散射、传输)、区块链网络选择和交易费用率。
  • 在具有波动信道状态、物联网设备状态和区块链动态的随机环境中,最大化网络吞吐量并最小化存储成本。

提出的方法

  • 为次级网关建立随机优化问题,联合决策在动态主信道状态、不确定物联网设备状态和不可预测区块链环境下的时间调度、区块链网络和交易费用率。
  • 采用深度强化学习(DRL)结合双值双重深度Q网络(D3QN),学习一种在吞吐量最大化与成本最小化之间取得平衡的最优策略。
  • 将系统建模为马尔可夫决策过程(MDP),其状态空间和动作空间较大,状态包括信道占用情况、设备能量和数据可用性以及区块链网络状况。
  • 设计综合奖励函数,整合网络吞吐量、交易成功率概率和成本效率,以指导D3QN的训练。
  • 采用基于仿真的训练与评估方法,将DRL策略与Q-learning、随机、HTT和背向散射策略等基线方法进行对比。
  • 在不同帧长、数据到达率、攻击者概率和可用区块链网络数量的条件下评估性能。

实验结果

研究问题

  • RQ1在动态且不确定的条件下,如何在基于区块链的射频供电背向散射认知无线电网络中联合优化时间调度、区块链网络选择和交易费用率?
  • RQ2帧长对所提框架中网络吞吐量和交易费用率有何影响?
  • RQ3与传统基线方案相比,数据到达率如何影响DRL调度策略的性能?
  • RQ4攻击者找到新区块的概率增加时,对不同区块链环境中吞吐量和交易费用率有何影响?
  • RQ5随着可用区块链网络数量的增加,吞吐量、交易费用和整体系统奖励将如何变化?

主要发现

  • 所提出的DRL方案在数据到达率为8时,平均吞吐量达到每帧9.1个数据单位,显著优于随机、HTT和背向散射方案(分别为4.4、3.3和1.6个单位)。
  • 由于在固定繁忙信道周期内背向散射操作保持一致,吞吐量随帧长增加而保持稳定。
  • 随着帧长增加,所有方案的交易费用率均因数据包尺寸增大而上升,但DRL方案仍保持成本效率。
  • 当攻击者找到新区块的概率增加时,所有方案的吞吐量均下降,交易费用急剧上升,但DRL方案维持最高性能。
  • 随着可用区块链网络数量的增加,系统性能得到提升:吞吐量增加,交易费用降低,总奖励上升,得益于更优的网络选择机会。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。