[论文解读] Wireless Edge-Empowered Metaverse: A Learning-Based Incentive Mechanism for Virtual Reality
本文提出一种基于深度强化学习(DRL)的双人荷兰式拍卖机制,用于优化无线边缘赋能的元宇宙中VR服务的分配与定价。通过将用户体验建模为感知质量(SSIM/VMAF),该框架实现了VR用户与服务提供商之间高效、低成本的交易,相较于基线方法,拍卖信息交换成本降低至少50%,同时实现接近最优的社会福利。
The Metaverse is regarded as the next-generation Internet paradigm that allows humans to play, work, and socialize in an alternative virtual world with immersive experience, for instance, via head-mounted display for Virtual Reality (VR) rendering. With the help of ubiquitous wireless connections and powerful edge computing technologies, VR users in wireless edge-empowered Metaverse can immerse in the virtual through the access of VR services offered by different providers. However, VR applications are computation- and communication-intensive. The VR service providers (SPs) have to optimize the VR service delivery efficiently and economically given their limited communication and computation resources. An incentive mechanism can be thus applied as an effective tool for managing VR services between providers and users. Therefore, in this paper, we propose a learning-based Incentive Mechanism framework for VR services in the Metaverse. First, we propose the quality of perception as the metric for VR users immersing in the virtual world. Second, for quick trading of VR services between VR users (i.e., buyers) and VR SPs (i.e., sellers), we design a double Dutch auction mechanism to determine optimal pricing and allocation rules in this market. Third, for auction communication reduction, we design a deep reinforcement learning-based auctioneer to accelerate this auction process. Experimental results demonstrate that the proposed framework can achieve near-optimal social welfare while reducing at least half of the auction information exchange cost than baseline methods.
研究动机与目标
- 设计一种激励机制,以高效管理无线边缘赋能的元宇宙中用户与服务提供商之间的VR服务交易。
- 通过联合使用SSIM和VMAF作为度量标准,重新定义VR中的效用,聚焦于感知质量。
- 通过叫市机制实现VR用户(买家)与服务提供商(卖家)之间的快速、异步交易。
- 通过基于DRL的拍卖师学习最优时钟更新策略,无需先验知识,从而降低拍卖信息交换成本。
- 在动态、大规模的VR服务市场中,平衡社会福利与拍卖效率。
提出的方法
- 提出一种基于感知的效用度量方法,利用结构相似性(SSIM)和视频多方法评估融合(VMAF)来量化VR用户的体验质量。
- 设计一种双人荷兰式拍卖(DDA)机制,采用递减的买家时钟和递增的卖家时钟,实现异步、有限时间内的交易。
- 引入一个深度强化学习(DRL)智能体作为拍卖师,根据市场反馈动态调整时钟步长。
- 使用一种奖励函数训练DRL智能体,以在最大化社会福利的同时最小化信息交换成本。
- 采用马尔可夫决策过程(MDP)框架,使拍卖师通过与VR服务市场参与者的交互学习最优策略。
- 设定拍卖参数:价格范围[1, 100],最小间隔1,权重w₁ = w₂ = 0.5,交换成本c = 0.1。

实验结果
研究问题
- RQ1如何有效建模非全景VR中的感知质量,以反映元宇宙中用户的实际体验?
- RQ2何种拍卖机制能够在叫市设置下实现VR用户与服务提供商之间高效、异步的匹配与定价?
- RQ3在动态的VR服务市场中,基于DRL的拍卖师是否能显著降低信息交换成本,同时不明显损害社会福利?
- RQ4基于DRL的拍卖师性能如何随市场规模和VR服务比特率需求的增加而变化?
- RQ5在无线边缘赋能的元宇宙应用中,激励机制下的社会福利与拍卖效率之间存在何种权衡?
主要发现
- 基于DRL的DDA实现了接近最优的社会福利,与最大可实现社会福利相比仅损失约5%。
- 所提出的框架相较于最先进方法将拍卖信息交换成本降低至少50%,相较于原始DDA基线降低三分之二。
- 在大规模市场(如50×50)中,基于DRL的DDA约在160个周期内收敛,但收敛速度较慢且稳定性低于小规模市场。
- 随着VR服务比特率提高,感知质量改善,社会福利随之提升,且基于DRL的DDA能动态适应此变化。
- 与随机方法相比,基于DRL的拍卖师实现了20%更高的社会福利,证明了学习策略相较于随机或固定策略的价值。
- 在大规模市场中,基于DRL的DDA仍保持高效率,其信息交换成本显著低于SOTA和原始DDA方法。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。