[Paper Review] Goal-oriented Semantic Communications for Avatar-centric Augmented Reality
This paper proposes a task-oriented and semantics-aware communication framework (TSAR) for avatar-centric augmented reality, leveraging semantic abstraction and shared base knowledge to reduce transmission latency by 95.6% and improve geometric and color fidelity by up to 82.4% and 20.4%, respectively, compared to traditional point cloud transmission in 6G wireless AR systems.
Upon the advent of the emerging metaverse and its related applications in Augmented Reality (AR), the current bit-oriented network struggles to support real-time changes for the vast amount of associated information, hindering its development. Thus, a critical revolution in the Sixth Generation (6G) networks is envisioned through the joint exploitation of information context and its importance to the task, leading to a communication paradigm shift towards semantic and effectiveness levels. However, current research has not yet proposed any explicit and systematic communication framework for AR applications that incorporate these two levels. To fill this research gap, this paper presents a task-oriented and semantics-aware communication framework for augmented reality (TSAR) to enhance communication efficiency and effectiveness in 6G. Specifically, we first analyse the traditional wireless AR point cloud communication framework and then summarize our proposed semantic information along with the end-to-end wireless communication. We then detail the design blocks of the TSAR framework, covering both semantic and effectiveness levels. Finally, numerous experiments have been conducted to demonstrate that, compared to the traditional point cloud communication framework, our proposed TSAR significantly reduces wireless AR application transmission latency by 95.6%, while improving communication effectiveness in geometry and color aspects by up to 82.4% and 20.4%, respectively.
Motivation & Objective
- Address the high bandwidth and low-latency demands of real-time avatar-centric augmented reality (AR) applications in the metaverse.
- Overcome limitations of traditional bit-oriented wireless communication in handling complex, dynamic AR data such as point clouds and avatars.
- Develop a systematic, end-to-end communication framework that integrates both semantic and effectiveness levels for AR applications.
- Improve communication efficiency and quality of experience (QoE) by focusing on task-relevant semantic content and structural relationships in avatar data.
Proposed method
- Design a graph-based representation to model relationships between semantic components of avatars, including geometry, color, and skeletal structure.
- Introduce a semantic information extractor using deep learning to identify and transmit only task-critical semantic features from the avatar data.
- Leverage shared base knowledge—such as pre-defined avatar skeletons and models—across transmitter and receiver to reduce redundant data transmission.
- Implement an end-to-end wireless communication pipeline where semantic features are extracted, compressed, and transmitted, then reconstructed at the receiver using shared knowledge.
- Use effectiveness-level metrics such as P2Point (point-to-point error) and PSNR_y (luminance PSNR) to evaluate geometric and color fidelity during reconstruction.
- Optimize transmission by replacing full point cloud data with compact skeletal representations (25 points) for pose updates, drastically reducing data volume.

Experimental results
Research questions
- RQ1How can semantic communication be systematically integrated into avatar-centric AR applications to improve efficiency and effectiveness?
- RQ2What role does shared base knowledge play in reducing transmission overhead while preserving avatar fidelity in wireless AR systems?
- RQ3To what extent can task-oriented semantic abstraction reduce latency and bandwidth consumption in AR point cloud transmission compared to traditional methods?
- RQ4How do semantic and effectiveness-level features jointly contribute to improved reconstruction quality in AR avatar rendering?
- RQ5What is the impact of wireless channel conditions (e.g., SNR) on the performance of semantic-aware AR communication frameworks?
Key findings
- The TSAR framework reduces transmission latency by 95.6% compared to traditional point cloud communication, significantly enhancing real-time performance.
- TSAR improves geometric fidelity (P2Point) by up to 82.4% and color fidelity (PSNR_y) by up to 20.4% under optimal SNR conditions.
- EC-TSAR and E-TSAR frameworks outperform TSAR and point cloud in both P2Point and PSNR_y metrics due to enhanced base knowledge utilization.
- Even at low SNR (below 8 dB), EC-TSAR and E-TSAR maintain stable performance with minimal distortion, while TSAR and point cloud show increasing distortion.
- The semantic extraction step adds only ~1 second per 100 frames, a negligible overhead compared to the massive reduction in transmitted data volume.
- Using 25 skeletal points instead of 2,048 point cloud points for pose updates reduces client-side rendering time and bandwidth usage substantially.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.