[Paper Review] Transient Neural Radiance Fields for Lidar View Synthesis and 3D Reconstruction
This paper proposes Transient Neural Radiance Fields (Transient NeRF), a novel method for 3D reconstruction and view synthesis from single-photon lidar data using transient photon arrival histograms. By modeling time-of-flight distributions as a neural radiance field, the approach enables high-fidelity 3D reconstruction and novel view synthesis from sparse, low-SNR lidar measurements, achieving state-of-the-art performance on a new 360° captured dataset with 6 real-world scenes.
Neural radiance fields (NeRFs) have become a ubiquitous tool for modeling scene appearance and geometry from multiview imagery. Recent work has also begun to explore how to use additional supervision from lidar or depth sensor measurements in the NeRF framework. However, previous lidar-supervised NeRFs focus on rendering conventional camera imagery and use lidar-derived point cloud data as auxiliary supervision; thus, they fail to incorporate the underlying image formation model of the lidar. Here, we propose a novel method for rendering transient NeRFs that take as input the raw, time-resolved photon count histograms measured by a single-photon lidar system, and we seek to render such histograms from novel views. Different from conventional NeRFs, the approach relies on a time-resolved version of the volume rendering equation to render the lidar measurements and capture transient light transport phenomena at picosecond timescales. We evaluate our method on a first-of-its-kind dataset of simulated and captured transient multiview scans from a prototype single-photon lidar. Overall, our work brings NeRFs to a new dimension of imaging at transient timescales, newly enabling rendering of transient imagery from novel views. Additionally, we show that our approach recovers improved geometry and conventional appearance compared to point cloud-based supervision when training on few input viewpoints. Transient NeRFs may be especially useful for applications which seek to simulate raw lidar measurements for downstream tasks in autonomous driving, robotics, and remote sensing.
Motivation & Objective
- Address the challenge of reconstructing high-quality 3D scenes from sparse, low-signal single-photon lidar measurements with limited viewpoints.
- Enable accurate novel view synthesis and depth estimation from transient photon arrival data using neural rendering.
- Develop a new 360° captured dataset with six real-world scenes scanned using a custom-built single-photon lidar system.
- Improve reconstruction fidelity and view synthesis quality by modeling time-of-flight distributions as a neural radiance field.
- Demonstrate the method's robustness across diverse materials, textures, and geometries using a calibrated hardware prototype.
Proposed method
- Construct a custom single-photon lidar system using a pulsed laser (532 nm, 35 ps pulses), 2D scanning mirrors, and a SPAD with TCSPC to capture photon arrival timestamps at 8 ps resolution over 1500 bins.
- Acquire 360° views of six real-world scenes by rotating each scene in 18° increments on a rotation stage, with 20 minutes of total exposure per scene.
- Calibrate camera intrinsics using the raxel model and a checkerboard on a translation stage, with subpixel corner refinement and polynomial fitting to correct for optical distortion.
- Estimate extrinsic parameters (center and axis of rotation) via a two-step procedure: coarse plane fitting on point clouds followed by optimization of 3D checkerboard corner alignment using Rodrigues' rotation formula.
- Train a neural radiance field that predicts volume density and view-dependent color from transient photon histograms, incorporating time-of-flight information as input features.
- Optimize the network using a differentiable rendering loss that minimizes L1 loss on depth and perceptual losses (LPIPS, SSIM) for novel view synthesis.
Experimental results
Research questions
- RQ1Can transient photon histograms from single-photon lidar be effectively modeled as a neural radiance field to enable high-fidelity 3D reconstruction and view synthesis?
- RQ2How does the proposed Transient NeRF method compare to existing NeRF-based approaches in terms of view synthesis quality and depth accuracy on real-world lidar data?
- RQ3What is the impact of using transient time-of-flight data as input to a neural radiance field for reconstruction from sparse and low-SNR measurements?
- RQ4How robust is the method across diverse scene geometries, materials, and textures under limited viewpoint sampling?
- RQ5To what extent can the calibrated hardware system and captured dataset support generalization and benchmarking in transient lidar reconstruction?
Key findings
- The proposed Transient NeRF achieves state-of-the-art performance on a new 360° captured dataset, with a mean LPIPS of 0.172 (5-view) and SSIM of 0.872 (5-view), significantly outperforming Instant NGP, DS-NeRF, and Urban NeRF baselines.
- For depth estimation, the method achieves a mean L1 depth error of 0.006 m (5 views), a 45% improvement over the next-best baseline (0.014 m) on the same dataset.
- On the 'boots' scene, the method achieves a 5-view SSIM of 0.914 and LPIPS of 0.155, demonstrating strong performance on challenging textures and specular surfaces.
- The method reduces L1 depth error by over 50% compared to the second-best baseline (Urban NeRF w/Mask) when using only 2 views, indicating strong generalization from sparse input.
- The calibration pipeline successfully estimates the center and axis of rotation with sub-millimeter accuracy, enabling precise extrinsic alignment across all 360° views.
- The dataset, captured with 1500 time bins at 8 ps resolution and 18° view increments, enables high-fidelity reconstruction and serves as a benchmark for future transient lidar research.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.