Skip to main content
QUICK REVIEW

[논문 리뷰] From Darkness to Detail: Frequency-Aware SSMs for Low-Light Vision

Eashan Adhikarla, Kai Zhang|arXiv (Cornell University)|2024. 08. 19.
Image Enhancement Techniques인용 수 4
한 줄 요약

ExpoMamba는 혼합 노출을 관리하고 텍스처 세부 정보를 향상시키기 위해 주파수 인식 Mamba 기반 아키텍처를 제안한다. 새로운 주파수 상태공간 블록(FSSB)을 활용하여 효율적이고 효과적인 저조도 영상 강화를 실현한다. 기존 모델 대비 2-3배 빠른 추론(36.6 ms)과 최대 20% 높은 PSNR 성능을 달성하여 실시간 엣지 배포에 적합하다.

ABSTRACT

Low-light image enhancement remains a persistent challenge in computer vision, where state-of-the-art models are often hampered by hardware constraints and computational inefficiency, particularly at high resolutions. While foundational architectures like transformers and diffusion models have advanced the field, their computational complexity limits their deployment on edge devices. We introduce ExpoMamba, a novel architecture that integrates a frequency-aware state-space model within a modified U-Net. ExpoMamba is designed to address mixed-exposure challenges by decoupling the modeling of amplitude (intensity) and phase (structure) in the frequency domain. This allows for targeted enhancement, making it highly effective for real-time applications, including downstream tasks like object detection and segmentation. Our experiments on six benchmark datasets show that ExpoMamba is up to 2-3x faster than competing models and achieves a 6.8\% PSNR improvement, establishing a new state-of-the-art in efficient, high-quality low-light enhancement. Source code: https://www.github.com/eashanadhikarla/ExpoMamba.

연구 동기 및 목표

  • 저조도 영상에서 과노출 및 과소노출 영역이 동시에 존재하는 혼합 노출 문제를 해결한다.
  • 실시간 엣지 응용에서 트랜스포머 및 디퓨전 기반 모델의 계산 비효율성과 높은 추론 지연 문제를 해결한다.
  • 저조도 강화에서 높은 정성적 및 구조적 영상 품질을 유지하는 경량이고 효율적인 아키텍처를 개발한다.
  • 다양한 해상도에서의 강력한 다중 해상도 추론을 가능하게 하기 위해 동적 배치 학습 기반 및 진폭과 위상 성분의 적응적 처리를 구현한다.

제안 방법

  • 장거리 의존성을 효율적으로 모델링하기 위해 수정된 U-Net에 2D-Mamba 블록과 주파수 도메인 처리를 통합한다.
  • 특징 맵을 진폭과 위상 성분으로 분해하여 대상 강화를 위한 주파수 상태공간 블록(FSSB)을 제안한다.
  • 진폭(노이즈/왜곡 억제용)과 위상(부드럽게 만들기 및 세부 정보 복구용) 성분을 별도의 2D-Mamba 블록으로 처리한다.
  • 주파수 도메인 처리 후 역 푸리에 변환을 통해 공간 도메인 출력을 재구성한다.
  • 다양한 입력 해상도와 조명 조건에서의 일반화를 향상시키기 위해 동적 배치 학습 기반을 구현한다.
  • 구조 유지 및 잡음 제거를 위해 L1, VGG, SSIM, LPIPS 및 과노출 정규화 항을 포함한 복합 손실 함수를 적용한다.
Figure 1: [ top: 400x600; bottom: 3840x2160] Scatter plot of model inference time vs. PSNR. Baselines that used ground-truth mean information to produce metrics were reproduced without such information for fairness.
Figure 1: [ top: 400x600; bottom: 3840x2160] Scatter plot of model inference time vs. PSNR. Baselines that used ground-truth mean information to produce metrics were reproduced without such information for fairness.

실험 결과

연구 질문

  • RQ1주파수 인식 상태공간 모델링이 계산 효율성을 유지하면서 저조도 영상 강화에 기여할 수 있는가?
  • RQ2주파수 도메인에서 진폭과 위상 처리를 분리함으로써 텍스처 복구 및 노이즈 억제에 어떤 기여를 하는가?
  • RQ3혼합 노출 상황에서 FSSB 블록이 표준 합성곱 및 어텐션 메커니즘보다 얼마나 뛰어난 성능을 보이는가?
  • RQ4입력 평균 밝기에 기반한 동적 추론 조정이 다양한 조명 조건에서 성능 향상에 기여하는가?
  • RQ5HDR 인식 구성요소(HDR, HDR-CSRNet+)의 통합이 저조도 강화에서 과노출 잡음 제거에 실제로 기여하는가?

주요 결과

  • ExpoMamba는 LOL-v1 데이터셋에서 PSNR 25.640을 달성하여 경쟁 모델 대비 15–20% 향상된 성능을 보였다.
  • 모델은 추론 시간을 36.6 ms로 줄여 기존 기준 대비 2–3배 빠른 처리 속도를 확보하여 실시간 응용에 적합하다.
  • FSSB 블록은 성능 향상에 기여하며, 추론 시 FSSB를 포함한 경우와 포함하지 않은 경우의 아블레이션 분석 결과, PSNR가 2.66 dB 향상되었다.
  • 추론 중 동적 조정(DA)을 적용하면 PSNR가 0.53 dB 향상되고 SSIM이 0.015 향상되어 입력 밝기 변화에 대한 강건성을 입증했다.
  • HDR-CSRNet+는 표준 HDR 및 HDROut 레이어보다 뛰어나지 않아, 아블레이션 연구에서 최고의 PSNR(25.640)와 SSIM(0.860) 성능을 기록했다.
  • L1, VGG, SSIM, LPIPS 및 과노출 정규화 손실 함수의 조합은 정성적 품질 향상과 잡음 제거에 뛰어난 효과를 보였다.
Figure 2: Overview of the ExpoMamba Architecture. The diagram illustrates the information flow through the ExpoMamba model. The architecture efficiently processes sRGB images by integrating convolutional layers, 2D-Mamba blocks, and deep supervision mechanisms to enhance image reconstruction, partic
Figure 2: Overview of the ExpoMamba Architecture. The diagram illustrates the information flow through the ExpoMamba model. The architecture efficiently processes sRGB images by integrating convolutional layers, 2D-Mamba blocks, and deep supervision mechanisms to enhance image reconstruction, partic

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.