[Paper Review] Rapid Whole Slide Imaging via Learning-based Two-shot Virtual Autofocusing
This paper proposes a learning-based two-shot virtual autofocusing (TSVA) method for rapid whole slide imaging (WSI), eliminating mechanical refocusing by recovering in-focus images from just two out-of-focus shots using a U-Net-inspired deep network. The method achieves high-throughput WSI with minimal scanning time and maintains high image quality, as shown by a 3.49 dB average PSNR gain over U-Net and only 0.12 average cell counting error on recovered images.
Whole slide imaging (WSI) is an emerging technology for digital pathology. The process of autofocusing is the main influence of the performance of WSI. Traditional autofocusing methods either are time-consuming due to repetitive mechanical motions, or require additional hardware and thus are not compatible to current WSI systems. In this paper, we propose the concept of extit{virtual autofocusing}, which does not rely on mechanical adjustment to conduct refocusing but instead recovers in-focus images in an offline learning-based manner. With the initial focal position, we only perform two-shot imaging, in contrast traditional methods commonly need to conduct as many as 21 times image shooting in each tile scanning. Considering that the two captured out-of-focus images retain pieces of partial information about the underlying in-focus image, we propose a U-Net-inspired deep neural network based approach for fusing them into a recovered in-focus image. The proposed scheme is fast in tissue slides scanning, enabling a high-throughput generation of digital pathology images. Experimental results demonstrate that our scheme achieves satisfactory refocusing performance.
Motivation & Objective
- To address the time-consuming nature of traditional WSI autofocusing methods that rely on repetitive mechanical z-stacking.
- To eliminate the need for high-precision mechanical systems by replacing in-scan autofocusing with offline learning-based image recovery.
- To enable high-throughput WSI by reducing image acquisition per tile from 21 shots to just two.
- To maintain high image quality and accuracy in downstream analysis despite reduced acquisition and no mechanical adjustment.
Proposed method
- The method uses a U-Net-inspired deep neural network (TSVA) to fuse two out-of-focus images captured at symmetric defocus offsets around the initial focal plane.
- The initial focal position is determined from a full z-stack on the first tile, which is then used as a reference for subsequent tiles.
- For each subsequent tile, only two images are captured—one at a defocus offset below and one above the estimated focal plane—reducing scanning time significantly.
- The two captured out-of-focus images are fed into the TSVA network to reconstruct a high-quality in-focus image in an offline, learning-based manner.
- The network is trained on large-scale paired data of out-of-focus and corresponding in-focus images to learn the mapping from defocused to focused appearance.
- The approach is compatible with existing WSI platforms as it does not require additional hardware beyond standard two-shot imaging.
Experimental results
Research questions
- RQ1Can a two-shot imaging strategy with learning-based recovery achieve comparable or better autofocusing performance than traditional multi-shot z-stack methods?
- RQ2How does the performance of the proposed virtual autofocusing method vary with different defocus offsets between the two captured images?
- RQ3To what extent does the use of dual input images improve the quality of the recovered in-focus image compared to single-image input?
- RQ4How robust is the proposed method across diverse tissue slide types and imaging conditions?
- RQ5Can the recovered in-focus images maintain sufficient quality for accurate downstream analysis such as cell counting?
Key findings
- The proposed TSVA method achieves an average PSNR gain of 3.49 dB over standard U-Net when using dual inputs, demonstrating superior image reconstruction quality.
- The highest PSNR gains occur at ±0.5 μm defocus offsets, which are also the most common estimated focal positions in practice.
- The method maintains strong generalization, achieving consistent performance across different test sets from diverse sources, with a 3.49 dB average PSNR gain over U-Net.
- The average cell counting error on recovered in-focus images is only 0.12, indicating minimal impact on downstream diagnostic accuracy.
- The method reduces image acquisition per tile from 21 shots to just 2, significantly accelerating scanning while maintaining high image quality.
- Subjective evaluations confirm that TSVA produces fewer structural errors than U-Net, especially in the critical ±1 μm defocus range.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.