Skip to main content
QUICK REVIEW

[논문 리뷰] Auditory Machine Learning Training and Testing Pipeline: AMLTTP v3.0

Ivo Trowitzsch, Youssef Kashef|arXiv (Cornell University)|2019. 02. 21.
Music and Audio Processing참고 문헌 8인용 수 12
한 줄 요약

이 논문은 14개의 클래스(예: 경보음, 개 짖는 소리, 피아노 등)에 속하는 고해상도 고음질 소리 이벤트 1,017건과 일반 소리 303건을 포함한 공개된 오디오 데이터베이스인 NIGENS를 소개한다. 모든 소리는 정확한 시작 및 종료 시간 태그가 부여되어 있으며, 음향 기계학습 모델의 강력한 훈련 및 테스트를 가능하게 하며, 특히 소리 이벤트 검출에 유용하다. AMLTTP v3.0을 통해 파이프라인 통합과 복잡한 시나리오 합성도 지원된다.

ABSTRACT

NIGENS (<strong>N</strong>eural <em><strong>I</strong></em>nformation Processing group <em><strong>GEN</strong></em>eral sounds) is a database provided for sound-related modeling in the field of computational auditory scene analysis, particularly for sound event detection, that has emerged from the Two!Ears project. It contains 1017 wav files of various lengths (between 1s and 5mins), in total comprising 4h:46m of sound material. Mostly, sounds are provided with 32-bit precision and 44100 Hz sampling rate. The files contain sound events in isolation, i.e. without superposition of ambient or other foreground sources. Fourteen distinct sound classes are included: <em>alarm</em>, <em>crying baby</em>, <em>crash</em>, <em>barking dog</em>, <em>running engine</em>, <em>burning fire</em>, <em>footsteps</em>, <em>knocking on door</em>, <em>female</em> and <em>male speech</em>, <em>female</em> and <em>male scream</em>, <em>ringing phone</em>, <em>piano</em>. Additionally, there is the <em>general</em> (“anything else”) class. Care has been taken to select sound classes representing different features, like noise-like or pronounced, discrete or continuous. The <em>general</em> class is a pool of sound events different than the 14 distuingished target sound classes, containing as heterogeneous sounds as possible (303 in total). For example, it includes nature sounds such as wind, rain, or animals, sounds from human-made environments such as honks, doors, or guns, as well as human sounds like coughs. These sounds are intended both as ``disturbance'' sound events (superposing) and as counterexamples to target sound classes.<br> <br> Wav files are accompanied by annotation (.txt) files that include perceptual on- and offset times of the file's sound events. You are free to use this database non-commercially under Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 license. <strong>If you use this data set, please cite as:</strong> <strong>Ivo Trowitzsch, Jalil Taghia, Youssef Kashef, and Klaus Obermayer (2019). <em>The NIGENS general sound events database</em>. Technische Universität Berlin, Tech. Rep. arXiv:1902.08314 [cs.SD] </strong> In [1], we have developed and analyzed a robust binaural sound event detection training scheme using NIGENS. In [2], we have extended it to join sound event detection and localization through spatial segregation. [1] Trowitzsch, I., Mohr, J., Kashef, Y., Obermayer, K. (2017). <em>Robust detection of environmental sounds in binaural auditory scenes</em>. IEEE/ACM Transactions on Audio, Speech, and Language Processing 25(6). [2] Trowitzsch, I., Schymura, C., Kolossa, D., Obermayer, K. (2019). <em>Joining Sound Event Detection and Localization Through Spatial Segregation</em>. accepted for publication in IEEE/ACM Transactions on Audio, Speech, and Language Processing. DOI: 10.1109/TASLP.2019.2958408. E-Preprint: arXiv:1904.00055 [cs.SD].

연구 동기 및 목표

  • 일반 소리 이벤트 검출(SoED) 연구를 위한 고품질, 고립된 소리 이벤트 데이터베이스의 부족 문제를 해결하기 위해.
  • 모델 훈련 정확도 향상을 위해 청각적으로 일치하는 시작 및 종료 시간 태그를 정밀하게 제공하기 위해.
  • 실제 환경에서 발생할 수 있는 알려지지 않은 소리 이벤트를 시뮬레이션하기 위해 다양한 '일반' 클래스를 포함시키기 위해.
  • 알 수 없는 간섭 요소와 복잡한 음향 환경을 다룰 수 있는 강력한 SoED 모델 개발을 지원하기 위해.
  • 공개 가능하고 철저히 태그가 부여된, 명확한 라이선싱 조건을 갖춘 표준화된 벤치마크와 재현 가능한 연구를 가능하게 하기 위해.

제안 방법

  • 고품질 32비트, 44.1kHz 샘플링을 갖춘 714개의 고립된 사운드 파일을 StockMusic.com에서 확보하고, 303개의 일반 사운드 파일을 Freesound.org에서 확보한다.
  • 청각적으로 일치하는 시작 및 종료 시간을 명확히 표기하기 위해 엄격한 태깅 기준을 적용하며, 중간의 정적 부분은 제외하여 이벤트의 고립성을 확보한다.
  • 14개의 목표 클래스(예: 경보음, 아기 울음, 엔진 등)와 모든 비목표 사운드를 포함하는 '일반' 클래스로 데이터셋을 구성한다.
  • 오디오, 파일 목록, 태깅 정보를 모델 훈련을 위해 Auditory Machine Learning Training and Testing Pipeline(AMLTTP v3.0)를 사용하여 처리한다.
  • AMLTTP를 통해 고립된 이벤트를 공간적·시간적 제어로 조합하여 다성분 음향 시나리오를 생성한다.
  • 이중 귀 신호의 500ms 세그먼트에서 분류 정확도의 혼동 행렬을 사용해 모델 성능을 검증한다.

실험 결과

연구 질문

  • RQ1대규모 고품질 고립 소리 이벤트 데이터베이스를 어떻게 구성할 수 있을까? 이는 강력한 소리 이벤트 검출 연구를 지원하기 위해.
  • RQ2청각적으로 정확한 시작 및 종료 태그는 모델의 일반화 능력과 검출 성능에 얼마나 기여하는가?
  • RQ3다양한 '일반' 사운드 클래스를 포함시키면, 알 수 없는 또는 어휘에 없는 사운드 이벤트에 대한 모델의 강건성은 어떻게 향상되는가?
  • RQ4AMLTTP v3.0은 NIGENS를 사용해 다성분 태그가 부여된 복잡한 음향 시나리오 모델의 훈련 및 테스트에 효과적으로 통합할 수 있는가?
  • RQ5분류 혼동 패tern을 통해 보여지는 바와 같이, 목표 사운드 클래스 간의 음향 겹침 정도는 어느 정도인가?

주요 결과

  • NIGENS는 총 1,017개의 오디오 파일로 구성되어 있으며, 총 길이가 4시간 45분 12초이다. 이 중 714개는 14개의 클래스에 속하는 고립된 이벤트이고, 303개는 일반 사운드 파일이다.
  • 여성의 목소리 모델은 실제 여성의 목소리 세그먼트에서 99%의 검출률을 기록하여 높은 감도를 보였다.
  • 오분류 비율을 분석한 결과, 음향적 겹침이 뚜렷하게 드러났다. 예를 들어, 여성의 목소리 모델은 피아노 세그먼트의 10%를 잘못 분류했다.
  • '일반' 사운드 클래스는 넓은 대역폭과 비정형 스펙트럼을 보이며, 포함된 비목표 사운드의 다양성을 반영한다.
  • 지속성 스펙트럼은 시간적·주파수적 구조를 효과적으로 시각화하며, 경보음과 전화음 사이, 엔진과 화재 소리 사이의 유사성을 보여준다.
  • 비목표 '일반' 사운드의 포함은 실제 환경의 음향 복잡성과 알 수 없는 간섭 요소를 시뮬레이션함으로써 모델의 강건성을 향상시킨다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.