[论文解读] Multi-Player Multi-Armed Bandit Based Resource Allocation for D2D Communications
本文提出了一种多用户多臂老虎机(MP-MAB)框架,用于在部分信道状态信息(CSI)条件下进行设备到设备(D2D)通信中的功率与资源分配,采用UCB1和Exp3算法平衡探索与利用。结果表明,在5G网络环境下,MP-UCB1相较于kth-UCB1、DLF和基于Exp3的策略,能实现更高的总吞吐量与更好的公平性,同时确保蜂窝用户(CUs)的QoS要求。
Device-to-device (D2D) communications is expected to play a significant role in increasing the system capacity of the fifth generation (5G) wireless networks. To accomplish this, efficient power and resource allocation algorithms need to be devised for the D2D users. Since the D2D users are treated as secondary users, their interference to the cellular users (CUs) should not hamper the CU communications. Most of the prior works on D2D resource allocation assume full channel state information (CSI) at the base station (BS). However, the required channel gains for the D2D pairs may not be known. To acquire these in a fast fading channel requires extra power and control overhead. In this paper, we assume partial CSI and formulate the D2D power and resource allocation problem as a multi-armed bandit problem. We propose a power allocation scheme for the D2D users in which the BS allocates power to the D2D users if a certain signal-to-interference-plus-noise ratio (SINR) is maintained for the CUs. In a single player environment a D2D user selects a CU in every time slot by employing UCB1 algorithm. Since this resource allocation problem can also be considered as an adversarial bandit problem we have applied the exponential-weight algorithm for exploration and exploitation (Exp3) to solve it. In a multiple player environment, we extend UCB1 and Exp3 to multiple D2D users. We also propose two algorithms that are based on distributed learning algorithm with fairness (DLF) and kth-UCB1 algorithms in which the D2D users are ranked. Our simulation results show that our proposed algorithms are fair and achieve good performance.
研究动机与目标
- 为解决在部分CSI条件下D2D通信中的功率与资源分配挑战,其中由于高开销,获取完整信道状态信息不切实际。
- 确保蜂窝用户(CUs)的SINR维持在最低要求之上,同时D2D用户可机会性地复用其频谱。
- 设计分布式学习算法,以最小化用户间碰撞,并确保多个竞争同一资源块的D2D用户之间的公平性。
- 将单用户老虎机算法(UCB1、Exp3)扩展至具有优先级排序的多用户场景,实现可扩展且公平的D2D接入。
提出的方法
- 将D2D功率与资源分配问题建模为部分CSI条件下的多用户多臂老虎机(MP-MAB)问题。
- 为每个D2D用户应用UCB1算法,以选择能最大化预期吞吐量的蜂窝用户资源块,平衡探索与利用。
- 将UCB1和Exp3扩展至多用户场景,引入机制以减少碰撞并确保公平性。
- 提出两种新算法——DLF(带公平性的分布式学习)和kth-UCB1,其中用户按等级排序,并基于其等级学习选择动作,从而减少干扰。
- 对CUs施加最小SINR阈值(γ^tgt);仅当CUs的SINR高于该阈值时,才向D2D用户分配功率。
- 将瞬时奖励(吞吐量)归一化至区间[0,1],以统一标准化所有用户与信道的老虎机反馈。
实验结果
研究问题
- RQ1在部分CSI条件下,D2D用户如何高效地从蜂窝用户中选择资源块,同时维持CUs的最低服务质量(QoS)?
- RQ2用户排序与公平性机制对多用户D2D老虎机学习中碰撞减少与系统吞吐量的影响如何?
- RQ3在多用户D2D资源分配场景中,UCB1与基于Exp3的策略在公平性与累积吞吐量方面表现如何比较?
- RQ4对抗性老虎机算法(如Exp3)能否在保证QoS约束的D2D频谱共享中有效应用?
- RQ5kth-UCB1与DLF相较于标准UCB1,在多用户D2D场景中能将公平性提升多少?
主要发现
- 由于碰撞率更低且对最优CUs的选择频率更高,MP-UCB1在所有评估算法中实现了最高的总吞吐量。
- 公平性百分比(即每个D2D用户选择其最优CUs的时间比例)最高的是MP-UCB1,其次为DLF和kth-UCB1,而Exp3的公平性最低。
- 尽管碰撞率较高且公平性较低,kth-UCB1与DLF仍能实现高总吞吐量,因为用户基于其等级选择最优动作,从而减少了干扰。
- CUs的总吞吐量保持稳定,接近恒定平均值,波动主要由功率分配决策与碰撞引起。
- 由于D2D链路通信距离更短且频谱效率更高,D2D用户的总吞吐量超过CUs。
- 所提算法确保CUs的SINR始终高于阈值γ^tgt,保障了QoS并防止性能下降。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。