Skip to main content
QUICK REVIEW

[논문 리뷰] End-to-End Multi-Task Learning for Adjustable Joint Noise Reduction and Hearing Loss Compensation

Philippe Gonzalez, Vera Margrethe Frederiksen|arXiv (Cornell University)|2026. 03. 20.
Hearing Loss and Rehabilitation인용 수 0
한 줄 요약

저자들은 추론 시 독립적으로 조정 가능한 양으로 노이즈 감소와 난청 보상을 공동으로 수행하는 엔드-투-엔드 멀티태스크 DNN을 제안하며, 청력도 입력에 의해 개인화되고 미분가능한 청각 모델로 학습된다.

ABSTRACT

A multi-task learning framework is proposed for optimizing a single deep neural network (DNN) for joint noise reduction (NR) and hearing loss compensation (HLC). A distinct training objective is defined for each task, and the DNN predicts two time-frequency masks. During inference, the amounts of NR and HLC can be adjusted independently by exponentiating each mask before combining them. In contrast to recent approaches that rely on training an auditory-model emulator to define a differentiable training objective, we propose an auditory model that is inherently differentiable, thus allowing end-to-end optimization. The audiogram is provided as an input to the DNN, thereby enabling listener-specific personalization without the need for retraining. Results show that the proposed approach not only allows adjusting the amounts of NR and HLC individually, but also improves objective metrics compared to optimizing a single training objective. It also outperforms a cascade of two DNNs that were separately trained for NR and HLC, and shows competitive HLC performance compared to a traditional hearing-aid prescription. To the best of our knowledge, this is the first study that uses an auditory model to train a single DNN for both NR and HLC across a wide range of listener profiles.

연구 동기 및 목표

  • NR(노이즈 감소)와 HLC(난청 보상)를 함께 해결하는 단일 DNN을 개발한다.
  • 추론 시 NR 및 HLC의 독립적인 조정을 마스크 지수화를 통해 가능하게 한다.
  • 재학습 없이 청취자의 청력도(audiogram)를 포함시켜 처리를 개인화한다.
  • 에뮬레이터 기반 훈련 없이도 엔드-투-엔드 최적화를 가능하게 하기 위해 미분가능한 청각 모델을 사용한다.

제안 방법

  • DNN가 예측하는 두 개의 시간-주파수 마스크를 정의한다: 하나는 NR용, 하나는 HLC용이다.
  • NR와 HLC에 대해 각각 다른 목적을 사용하여 학습하고, 불확실성 기반 가중치 체계를 사용해 이를 균형 잡는다.
  • 추론 시 두 마스크를 독립적인 alphA_NR 및 alpha_HLC 매개변수로 각 마스크를 지수화하여 결합한다.
  • 모델 입력에는 청력도(audiogram)가 포함되어 청취자별 개인화를 가능하게 한다.
  • NR와 HLC 모두에 대해 생리학적으로 근거 있는, 학습 가능한 타깃을 제공하기 위해 미분가능한 청각 모델을 사용한다.

실험 결과

연구 질문

  • RQ1넓은 범위의 HI 청취자에 대해 NR과 HLC를 함께 수행하도록 단일 DNN을 엔드투엔드로 학습시킬 수 있는가?
  • RQ2각각의 목표를 가진 멀티태스크 학습이 단일 태스크나 계단식 접근 방식보다 객관적 지표를 향상시키는가?
  • RQ3재학습 없이 추론 시 NR과 HLC를 독립적으로 조정할 수 있는가?
  • RQ4입력으로 청력도(audiogram)를 포함시키는 것이 리스너별 재학습 없이도 효과적인 개인화를 가능하게 하는가?

주요 결과

  • 제안된 방법은 추론 시 지수화된 마스크를 통해 NR과 HLC의 독립적 조정을 가능하게 한다.
  • 불확실성 기반 가중치를 갖는 멀티태스크 학습이 단일 목표 최적화보다 객관적 지표를 향상시킨다.
  • NR과 HLC를 각각 독립적으로 학습한 두 DNN의 계단식 시스템을 능가한다.
  • 전통적인 보청 처방과 비교해도 경쟁력 있는 HLC 성능을 보이며 다양한 청취자 청력도에서 작동한다.
  • 저자들에 따르면, 다양한 청취자 프로필에 걸쳐 NR과 HLC 모두에 대해 단일 DNN을 학습시키기 위해 미분가능한 청각 모델을 사용하는 것은 이번이 처음이다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.