[论文解读] Dialog+ in Broadcasting: First Field Tests Using Deep-Learning-Based Dialogue Enhancement
本文提出 Dialog+,一种基于深度学习的系统,通过允许用户调节传统混音音频轨道中的对话音量,提升广播音频的语音可懂度,无需依赖基于对象的音频工作流程。对超过2,000名参与者的实地测试表明,90%的60岁以上观众经常或非常频繁地难以听清语音,83%的受访者更偏好使用 Dialog+ 的增强效果,证明其在提升主流观众可访问性方面的有效性。
Difficulties in following speech due to loud background sounds are common in broadcasting. Object-based audio, e.g., MPEG-H Audio solves this problem by providing a user-adjustable speech level. While object-based audio is gaining momentum, transitioning to it requires time and effort. Also, lots of content exists, produced and archived outside the object-based workflows. To address this, Fraunhofer IIS has developed a deep-learning solution called Dialog+, capable of enabling speech level personalization also for content with only the final audio tracks available. This paper reports on public field tests evaluating Dialog+, conducted together with Westdeutscher Rundfunk (WDR) and Bayerischer Rundfunk (BR), starting from September 2020. To our knowledge, these are the first large-scale tests of this kind. As part of one of these, a survey with more than 2,000 participants showed that 90% of the people above 60 years old have problems in understanding speech in TV "often" or "very often". Overall, 83% of the participants liked the possibility to switch to Dialog+, including those who do not normally struggle with speech intelligibility. Dialog+ introduces a clear benefit for the audience, filling the gap between object-based broadcasting and traditionally produced material.
研究动机与目标
- 解决因背景音量过大导致的广播电视语音可懂度问题,尤其针对年长或听力受损的观众。
- 为缺乏基于对象的音频元数据或工作流程的传统音频内容,实现个性化对话音量调节。
- 弥合现代基于对象的音频系统与传统制作的广播素材之间的差距。
- 在真实世界环境中,通过大规模实地测试,评估基于深度学习的对话增强解决方案的有效性与用户接受度。
提出的方法
- 训练深度神经网络,从混合的立体声或单声道音频轨道中分离并增强对话成分。
- 基于语音可懂度的感知模型,估算随时间变化的对话音量,可由用户按需调节。
- 在不需重新编码或更改元数据的情况下,实时应用于现有广播音频流。
- 将该解决方案集成至公共广播机构 WDR 和 BR 的实际广播工作流程中,用于实地测试。
- 通过涵盖不同年龄群体的真实观众大规模调查,评估用户偏好与可用性。
- 系统设计兼容现有分发基础设施,确保与标准音频格式的向后兼容性。
实验结果
研究问题
- RQ1基于深度学习的系统能否在不依赖基于对象的音频工作流程的前提下,有效提升标准广播音频中的对话可懂度?
- RQ2观众,尤其是年长成年人,在真实广播环境中如何感知并响应可调节的对话音量?
- RQ3与标准音频相比,Dialog+ 在提升听力困难观众的语音理解方面能带来多大程度的改善?
- RQ4Dialog+ 在不同人口统计群体中的用户接受率如何,包括那些通常不受语音可懂度问题影响的群体?
- RQ5后处理、非侵入式的增强系统能否实现与原生基于对象的音频解决方案相当的感知优势?
主要发现
- 90% 的60岁以上参与者表示在电视广播中经常或非常频繁地难以理解语音。
- 83% 的调查参与者,包括那些通常不受语音可懂度问题影响的群体,表达了对 Dialog+ 增强效果的偏好。
- 本次实地测试是首次在真实公共广播环境中对基于深度学习的对话增强系统进行的大规模实际评估。
- Dialog+ 有效实现了对非基于对象音频工作流程制作内容的用户可调节对话音量功能。
- 该系统展现出显著的感知优势和高用户接受度,证明其作为通往未来基于对象音频生态系统的有效桥梁的可行性。
- 结果证实,基于深度学习的增强技术可显著提升主流广播观众的可访问性与用户体验。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。