[论文解读] Soundscape Captioning using Sound Affective Quality Network and Large Language Model
本文提出了声音景观描述(SoundSCap)这一自动化任务,通过整合声音场景/事件识别与感知情感特质,生成上下文感知的音频场景描述。所提出的SoundSCaper模型结合了多尺度声学模型(SoundAQnet)与大语言模型(LLM),生成的描述在人类评估中达到专家水平质量,评分无统计学显著差异(满分5分下平均差异为0.21–0.25)。
We live in a rich and varied acoustic world, which is experienced by individuals or communities as a soundscape. Computational auditory scene analysis, disentangling acoustic scenes by detecting and classifying events, focuses on objective attributes of sounds, such as their category and temporal characteristics, ignoring their effects on people, such as the emotions they evoke within a context. To fill this gap, we propose the affective soundscape captioning (ASSC) task, which enables automated soundscape analysis, thus avoiding labour-intensive subjective ratings and surveys in conventional methods. With soundscape captioning, context-aware descriptions are generated for soundscape by capturing the acoustic scenes (ASs), audio events (AEs) information, and the corresponding human affective qualities (AQs). To this end, we propose an automatic soundscape captioner (SoundSCaper) system composed of an acoustic model, i.e. SoundAQnet, and a large language model (LLM). SoundAQnet simultaneously models multi-scale information about ASs, AEs, and perceived AQs, while the LLM describes the soundscape with captions by parsing the information captured with SoundAQnet. SoundSCaper is assessed by two juries of 32 people. In expert evaluation, the average score of SoundSCaper-generated captions is slightly lower than that of two soundscape experts on the evaluation set D1 and the external mixed dataset D2, but not statistically significant. In layperson evaluation, SoundSCaper outperforms soundscape experts in several metrics. In addition to human evaluation, compared to other automated audio captioning systems with and without LLM, SoundSCaper performs better on the ASSC task in several NLP-based metrics. Overall, SoundSCaper performs well in human subjective evaluation and various objective captioning metrics, and the generated captions are comparable to those annotated by soundscape experts.
研究动机与目标
- 自动化传统上依赖调查与专家评分的主观声音景观评估的繁重流程。
- 弥合客观音频事件检测与主观人类情感反应在声学环境中的差距。
- 开发统一框架,捕捉声音场景、事件及感知情感特质(PAQ),用于自然语言描述。
- 通过由16位声音景观专家组成的评审团,评估自动化描述系统与人类专家的性能对比。
- 利用深度学习与大语言模型(LLM),实现可扩展、上下文感知且情感敏感的复杂音频环境描述。
提出的方法
- 提出SoundAQnet,一种轻量化、多尺度深度神经网络,从原始音频中联合建模声音场景(AS)、音频事件(AE)与感知情感特质(PAQ)。
- 在大规模数据集上进行训练,结合参与者评分,学习如愉悦度、事件丰富度与平静度等感知属性,采用8D声音景观环形模型(SCM)进行表征。
- 利用多尺度表征捕捉短期显著事件(如警报声)与长期环境背景(如城市与森林环境)。
- 将SoundAQnet的输出与通用大语言模型(LLM)集成,从AS、AE与PAQ三个视角生成自然语言描述。
- 采用零样本或少样本提示策略,引导LLM从模型潜在表征中合成连贯、上下文感知的描述。
- 通过由评审团主导的人工评估,在保留测试集与外部混合数据集(涵盖不同音频长度与声学特性)上验证系统性能。
实验结果
研究问题
- RQ1自动化系统能否生成与专家标注描述质量相当的上下文感知声音景观描述?
- RQ2将声音场景/事件识别与感知情感特质(PAQ)结合,相较于传统AS/AE分类,能否显著提升描述性能?
- RQ3该模型在具有不同长度与声学特性的多样化音频片段上,泛化能力如何?
- RQ4SoundAQnet的情感特质预测与个体专家感知相比,在主观性与一致性方面表现如何?
- RQ5SoundSCaper生成的描述与专家标注描述之间的性能差距是否具有统计学显著性?
主要发现
- 在测试集上,SoundSCaper生成的描述平均人类评估得分为4.29(满分5分),专家标注描述为4.50,差异为0.21分。
- 在包含5个模型未知数据集的外部混合数据集上,SoundSCaper的平均得分为4.25,专家标注描述为4.50,差异为0.25分。
- SoundSCaper与专家标注描述之间的性能差距无统计学显著性,表明在人类感知中质量相当。
- SoundSCaper在捕捉主导音频事件与整体声音场景特征方面表现优异,但缺乏对空间与时间细节的精细刻画(如车辆方向、短时警报声存在)。
- SoundAQnet的情感特质预测比个体专家响应更趋于中性与概括,降低了个人偏见,但可能损失细微情感表达。
- 该模型在具有不同长度与声学特性的多样化音频片段上表现出强大泛化能力,证实其在真实环境中的鲁棒性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。