[Paper Review] FVV Live: A real-time free-viewpoint video system with consumer electronics hardware
FVV Live is a real-time, low-cost free-viewpoint video system using consumer-grade stereo cameras and off-the-shelf hardware, enabling high-quality virtual view synthesis through depth correction, lossless 12-bit depth coding, and layered view synthesis. It achieves sub-50ms motion-to-photon delay and end-to-end latency of ~250ms, with subjective assessments showing it significantly outperforms state-of-the-art DIBR methods.
© 2021 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.
Motivation & Objective
- To develop a low-cost, real-time free-viewpoint video system using only off-the-shelf components.
- To address the trade-off between video quality, real-time performance, and deployment cost in FVV systems.
- To enable seamless virtual navigation and bilateral immersive communication in applications like event broadcasting and videoconferencing.
- To improve synthesis quality despite limitations in consumer-grade depth estimation and bandwidth.
Proposed method
- Uses nine Stereolabs ZED stereo cameras arranged in a sparse array to capture multiview plus depth (MVD) data.
- Employs capture servers with depth post-processing to correct errors from stereo calibration inaccuracies.
- Applies lossless 12-bit depth coding over standard 8-bit video codecs to preserve high-precision depth data.
- Implements adaptive streaming that enables/disables camera streams based on virtual viewpoint position to reduce bitrate.
- Uses a layered view synthesis approach on a single edge server, separating foreground and background for efficient rendering.
- Performs real-time foreground/background segmentation in capture servers to enable layered rendering and optimize bandwidth.
Experimental results
Research questions
- RQ1Can a high-quality free-viewpoint video system be built using only consumer electronics hardware without sacrificing real-time performance?
- RQ2How can depth estimation errors from low-cost stereo cameras be effectively corrected in real time?
- RQ3What compression and transmission strategies maximize synthesis quality under limited bandwidth while preserving critical depth data?
- RQ4How does layered view synthesis with separate foreground/background treatment improve visual quality and reduce computational load?
- RQ5To what extent does FVV Live outperform state-of-the-art DIBR-based systems in subjective quality assessments?
Key findings
- FVV Live achieves a mean motion-to-photon delay of less than 50 ms and an end-to-end delay of approximately 250 ms, enabling real-time, responsive virtual navigation.
- Subjective evaluations show FVV Live is preferred over VSRS (a state-of-the-art DIBR method) in over 80% of comparisons across all scenarios, trajectories, and camera configurations.
- The system maintains consistent visual quality along virtual camera paths, with no statistically significant difference in quality between Still Camera 1/2 and 1/3 trajectories.
- DMOS scores for FVV Live are above 3.0 (fair quality) and reach up to 4.0–4.5 on smaller smartphone screens, indicating near-indistinguishable quality from physical camera references in simpler scenes.
- Lossless 12-bit depth coding significantly improves synthesis quality, and adaptive streaming reduces bandwidth by selectively transmitting only relevant camera data.
- The layered synthesis approach with real-time FG/BG segmentation enables efficient rendering and allows re-allocation of bandwidth to high-fidelity depth data.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.