Skip to main content
QUICK REVIEW

[论文解读] Task-Oriented Image Semantic Communication Based on Rate-Distortion Theory

Fangfang Liu, Wanjie Tong|arXiv (Cornell University)|Jan 26, 2022
AI in cancer detection被引用 5
一句话总结

本文提出了一种基于语义重建的任务导向语义通信框架(TOSC-SR),这是一种新颖的率失真框架,通过利用重建图像与任务标签之间的互信息,联合最小化像素级失真和与任务相关的语义失真。该方法在图像重建、人工智能任务准确率以及多任务泛化能力方面,均优于传统及基于深度学习的基准方法。

ABSTRACT

Task-oriented image semantic communication is a new communication paradigm, which aims to transmit semantics for artificial intelligent (AI) tasks while ignoring the reconstruction quality of the images. However, in some applications, such as autonomous driving, both image reconstruction quality and the performance of the followed AI tasks must be simultaneously considered. To tackle this challenge, this paper proposes a task-oriented semantic communication scheme with semantic reconstruction (TOSC-SR). Its main goal is to simultaneously minimize pixel-level and task-relevant semantic-level distortion during communications under a certain rate, which formulates a new rate-distortion optimization problem. To successfully measure the loss at the semantic level, a new form of semantic distortion measured by the mutual information between the semantic-reconstructed images and the task labels is proposed. Then, we derive an analytical solution for the formulated problem, where the self-consistent equations of the problem are obtained to determine the optimal mapping of the source and the semantic-reconstructed images. To implement TOSC-SR, we further obtain an extended form of rate-distortion form based on the variational approximation of mutual information, which is applicable to multiple AI tasks. Experimental results show that the proposed approach outperforms the traditional JPEG, JPEG2000, BPG, VVC-based image communication systems and deep learning based benchmarks in terms of image reconstruction quality, AI task performance, and multi-task generalization ability.

研究动机与目标

  • 为解决现有语义通信系统仅关注任务性能或图像重建的局限性。
  • 在固定传输速率下,同时优化像素级重建质量和与任务相关的语义保真度。
  • 基于信息论原理,提出一种同时包含两类失真类型的率失真优化问题的新形式。
  • 开发一种变分近似互信息方法,以实现可扩展的多任务语义通信。

提出的方法

  • 提出一种基于重建图像与任务标签之间互信息的新语义失真度量。
  • 通过源图像映射与重建图像映射的自洽方程,推导出联合率失真优化问题的解析解。
  • 引入互信息的变分近似方法,以支持端到端训练,并扩展至多种AI任务的应用。
  • 采用拉格朗日松弛框架,通过统一的优化目标,平衡速率、像素级失真与语义失真。
  • 采用参数化信道模型,其中条件概率 $ p(\hat{\bm{x}}|\bm{x}) $ 通过推导出的更新规则进行优化。
  • 引入正则化项 $ r(\bm{x}) $,用于吸收特定任务信息,从而实现对不同下游任务的灵活适应。

实验结果

研究问题

  • RQ1如何使图像通信系统同时优化图像重建质量与下游AI任务性能?
  • RQ2一种有效且基于信息论的语义失真度量应如何设计,以准确捕捉任务相关性?
  • RQ3在固定速率约束下,能否构建一个统一的率失真框架,联合最小化像素级与语义级失真?
  • RQ4在深度学习设置中,如何有效近似并优化重建图像与任务标签之间的互信息?
  • RQ5在多任务场景下,所提出方法相较于传统编解码器与现有语义通信系统,性能提升如何?

主要发现

  • 所提出的TOSC-SR框架在图像重建质量与下游AI任务性能方面,均优于JPEG、JPEG2000、BPG与VVC等传统编解码器。
  • 与最先进的基于深度学习的图像压缩方法相比,TOSC-SR在基准数据集上实现了更高的峰值信噪比(PSNR)与更好的分类准确率。
  • 该系统展现出强大的多任务泛化能力,在无需微调的情况下,可在多种不同AI任务中保持高性能。
  • 互信息的变分近似方法实现了多任务场景下有效且可扩展的优化。
  • 基于自洽方程推导出的解析解,为语义通信中最优映射设计提供了理论基础。
  • 实验结果证实,联合最小化像素级与语义级失真可显著提升整体系统性能。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。