Skip to main content
QUICK REVIEW

[Paper Review] iToF2dToF: A Robust and Flexible Representation for Data-Driven Time-of-Flight Imaging

Felipe Gutierrez, Huaijin Chen|arXiv (Cornell University)|Mar 11, 2021
Advanced Optical Sensing Technologies60 references32 citations
TL;DR

This paper proposes iToF2dToF, a data-driven method that reconstructs transient images from sparse iToF frequency measurements by interpolating/extrapolating frequencies, enabling robust depth estimation via rule-based peak finding. It improves accuracy under low SNR and challenging scenarios like specular MPI and optical cross-talk without retraining, outperforming prior methods on real-world and synthetic data with state-of-the-art results.

ABSTRACT

Indirect Time-of-Flight (iToF) cameras are a promising depth sensing technology. However, they are prone to errors caused by multi-path interference (MPI) and low signal-to-noise ratio (SNR). Traditional methods, after denoising, mitigate MPI by estimating a transient image that encodes depths. Recently, data-driven methods that jointly denoise and mitigate MPI have become state-of-the-art without using the intermediate transient representation. In this paper, we propose to revisit the transient representation. Using data-driven priors, we interpolate/extrapolate iToF frequencies and use them to estimate the transient image. Given direct ToF (dToF) sensors capture transient images, we name our method iToF2dToF. The transient representation is flexible. It can be integrated with different rule-based depth sensing algorithms that are robust to low SNR and can deal with ambiguous scenarios that arise in practice (e.g., specular MPI, optical cross-talk). We demonstrate the benefits of iToF2dToF over previous methods in real depth sensing scenarios.

Motivation & Objective

  • Address the limitations of data-driven iToF methods that lack explicit supervision for ambiguous scenarios like specular MPI and optical cross-talk.
  • Overcome the poor generalization of end-to-end models by introducing an intermediate transient image representation.
  • Enable robust depth estimation under low SNR conditions using a flexible, learnable frequency interpolation framework.
  • Integrate with simple rule-based peak-finding algorithms to maintain real-time performance and adaptability.
  • Demonstrate superior performance on real-world data and a new synthetic dataset with realistic noise and complex light transport.

Proposed method

  • Input sparse iToF frequency measurements (e.g., 20 MHz and 100 MHz) to a deep neural network for frequency interpolation/extrapolation.
  • Train the network to predict a dense set of iToF frequencies (e.g., 20–400 MHz), enabling high-resolution transient image reconstruction.
  • Use the reconstructed frequency data to estimate the direct-to-reflective (dToF) transient image via inverse Fourier transform.
  • Apply rule-based peak-finding algorithms (e.g., max-peak detection) to the dToF waveforms to extract depth estimates.
  • Leverage the dToF representation to separate direct and indirect light paths, mitigating multi-path interference (MPI) without requiring explicit supervision for each scenario.
  • Use a synthetic dataset with realistic geometry, textures, and noise to train and validate the model, ensuring generalization to real-world conditions.

Experimental results

Research questions

  • RQ1Can a data-driven model that reconstructs intermediate transient images from sparse iToF frequencies generalize better to ambiguous depth sensing scenarios than end-to-end models?
  • RQ2How does the iToF2dToF framework perform under low SNR conditions where traditional iToF and data-driven methods fail?
  • RQ3Can the dToF representation enable robust depth estimation in challenging scenarios such as specular reflections and optical cross-talk without retraining?
  • RQ4To what extent does the proposed frequency interpolation strategy improve transient image reconstruction accuracy compared to baseline methods?
  • RQ5Does integrating iToF2dToF with simple rule-based peak finding yield better depth accuracy than end-to-end learning on real-world data?

Key findings

  • iToF2dToF achieves state-of-the-art depth accuracy on real-world scenes with ground truth, outperforming both traditional phasor and data-driven iToF2Depth methods.
  • At 0.1ms exposure time (extremely low SNR), iToF2dToF maintains accurate depth estimation in regions where other methods fail or produce significant errors.
  • The method reduces median depth error by up to 50% compared to iToF2Depth on challenging scenes with specular reflections and optical cross-talk.
  • The synthetic dataset with realistic noise and complex light transport enables effective training and generalization to real-world scenarios.
  • iToF2dToF successfully reconstructs transient waveforms with minimal artifacts even from only two input frequencies, enabling high-quality depth maps.
  • The framework maintains robustness across varying fog densities, with accurate depth recovery even under high scattering conditions.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.