[Paper Review] Multi-focus Image Fusion: A Benchmark
This paper introduces the first comprehensive benchmark for multi-focus image fusion (MFIFB), comprising 105 image pairs, 30 algorithms, and 20 evaluation metrics. It reveals that deep learning-based methods underperform on challenging real-world datasets compared to conventional methods, primarily due to poor generalization from simulated training data lacking strong defocus spread effects.
Multi-focus image fusion (MFIF) has attracted considerable interests due to its numerous applications. While much progress has been made in recent years with efforts on developing various MFIF algorithms, some issues significantly hinder the fair and comprehensive performance comparison of MFIF methods, such as the lack of large-scale test set and the random choices of objective evaluation metrics in the literature. To solve these issues, this paper presents a multi-focus image fusion benchmark (MFIFB) which consists a test set of 105 image pairs, a code library of 30 MFIF algorithms, and 20 evaluation metrics. MFIFB is the first benchmark in the field of MFIF and provides the community a platform to compare MFIF algorithms fairly and comprehensively. Extensive experiments have been conducted using the proposed MFIFB to understand the performance of these algorithms. By analyzing the experimental results, effective MFIF algorithms are identified. More importantly, some observations on the status of the MFIF field are given, which can help to understand this field better.
Motivation & Objective
- Address the lack of a standardized benchmark for fair and comprehensive performance comparison in multi-focus image fusion (MFIF).
- Overcome the limitations of existing datasets like Lytro, which contain only 20 image pairs and lack strong defocus spread effects.
- Provide a unified platform with diverse algorithms, evaluation metrics, and a large-scale test set to enable objective performance evaluation.
- Identify the true state-of-the-art in MFIF by evaluating all integrated algorithms under consistent conditions.
- Highlight the gap between claimed performance and actual generalization ability of deep learning-based MFIF methods on real-world data.
Proposed method
- Construct a large-scale test dataset of 105 multi-focus image pairs, including Lytro, MFFW, and datasets from Savic et al. and Aymaz et al. to increase diversity and challenge.
- Collect and integrate 30 state-of-the-art MFIF algorithms, including spatial domain-based, transform domain-based, and deep learning-based methods.
- Implement 20 objective evaluation metrics spanning information theory (e.g., MI, NMI), structural similarity (SSIM, Q_Y), and quality assessment (PSNR, VIF, FMI).
- Conduct extensive quantitative and qualitative experiments using the benchmark to evaluate all algorithms under identical conditions.
- Use consistent training and inference protocols—no online model updates—ensuring fair comparison across deep learning and non-deep learning methods.
- Analyze performance across multiple subsets (e.g., Lytro, MFFW) to assess generalization and robustness of algorithms.
Experimental results
Research questions
- RQ1How do deep learning-based MFIF methods perform on a large-scale, diverse, and challenging real-world dataset compared to conventional methods?
- RQ2To what extent do current MFIF evaluation metrics provide consistent and reliable performance rankings across different algorithms?
- RQ3Why do many deep learning-based MFIF methods fail to generalize beyond datasets like Lytro, which lack strong defocus spread effects?
- RQ4Can a unified benchmark platform improve the fairness and comprehensiveness of MFIF algorithm evaluation compared to prior isolated studies?
- RQ5How consistent are qualitative visual results with quantitative metric rankings across different MFIF algorithms?
Key findings
- Deep learning-based MFIF methods, such as DPRL, rank only fifth on the full MFIFB dataset, despite being claimed as state-of-the-art in prior work.
- Conventional methods (spatial and transform domain-based) outperform deep learning-based methods on the MFIFB dataset, especially on subsets like MFFW with strong defocus spread effects.
- The performance gap arises because most deep learning models are trained on simulated data lacking realistic defocus spread, limiting their generalization to real-world scenarios.
- Evaluation results show that no single metric consistently ranks algorithms the same—different metrics favor different methods, highlighting the need for multi-metric evaluation.
- Qualitative and quantitative results often disagree, indicating that visual inspection and numerical metrics must both be used to evaluate MFIF performance.
- The MFIFB benchmark reveals that prior performance claims were often based on small datasets and biased metric selections, leading to overestimation of deep learning model capabilities.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.