[论文解读] Cross-Corpora Spoken Language Identification with Domain Diversification and Generalization
本文提出领域多样化与领域泛化技术,以提升低资源印度语在跨语料库语音语言识别(LID)中的表现。通过使用最大程度关注多样性的级联音频增强方法,并将增强类型视为伪领域,作者通过领域不变与领域感知学习增强ECAPA-TDNN模型,与基线相比,跨语料库EER最高降低5.23%。
This work addresses the cross-corpora generalization issue for the low-resourced spoken language identification (LID) problem. We have conducted the experiments in the context of Indian LID and identified strikingly poor cross-corpora generalization due to corpora-dependent non-lingual biases. Our contribution to this work is twofold. First, we propose domain diversification, which diversifies the limited training data using different audio data augmentation methods. We then propose the concept of maximally diversity-aware cascaded augmentations and optimize the augmentation fold-factor for effective diversification of the training data. Second, we introduce the idea of domain generalization considering the augmentation methods as pseudo-domains. Towards this, we investigate both domain-invariant and domain-aware approaches. Our LID system is based on the state-of-the-art emphasized channel attention, propagation, and aggregation based time delay neural network (ECAPA-TDNN) architecture. We have conducted extensive experiments with three widely used corpora for Indian LID research. In addition, we conduct a final blind evaluation of our proposed methods on the Indian subset of VoxLingua107 corpus collected in the wild. Our experiments demonstrate that the proposed domain diversification is more promising over commonly used simple augmentation methods. The study also reveals that domain generalization is a more effective solution than domain diversification. We also notice that domain-aware learning performs better for same-corpora LID, whereas domain-invariant learning is more suitable for cross-corpora generalization. Compared to basic ECAPA-TDNN, its proposed domain-invariant extensions improve the cross-corpora EER up to 5.23%. In contrast, the proposed domain-aware extensions also improve performance for same-corpora test scenarios.
研究动机与目标
- 解决因语料依赖的非语言偏差导致的低资源LID系统跨语料库泛化能力差的问题。
- 提升在有限内部语料上训练的LID模型在未见、多样化测试语料上的鲁棒性。
- 开发有效的数据增强策略,以模拟真实世界中的领域差异,增强训练数据多样性。
- 提出基于增强的伪领域进行领域泛化,以在无需目标领域数据的情况下提升泛化能力。
- 比较领域不变与领域感知学习策略在同语料库与跨语料库设置下的最优性能表现。
提出的方法
- 提出使用多种音频增强方法(如语音增强、编解码器类、噪声注入等)进行领域多样化,以增加训练数据的变异性。
- 引入最大多样性感知的级联增强方法,并通过优化折叠因子,在多样性与模型性能之间实现平衡。
- 将增强类型视为伪领域,以在缺乏真实目标领域数据的情况下实现领域泛化(DG)。
- 通过梯度反转与MMD-based损失实现领域不变学习,以最小化领域特定表示。
- 通过多任务学习实现领域感知学习,使模型能够同时学习语言与领域特定线索。
- 在ECAPA-TDNN上进行训练与评估,该模型是语音与语言识别的最先进深度神经网络架构,并扩展其用于DG。
实验结果
研究问题
- RQ1音频数据增强策略是否能有效模拟未见的领域差异,从而提升跨语料库LID的泛化能力?
- RQ2将增强类型视为伪领域是否能在低资源LID中实现有效的领域泛化?
- RQ3在跨语料库LID泛化中,领域不变学习是否比领域感知学习更有效?
- RQ4与标准增强方法相比,领域多样化在未见语料的盲评估中表现如何?
- RQ5领域泛化技术是否能缩小同语料库与跨语料库评估之间的性能差距?
主要发现
- 所提出的最大多样性感知级联增强的领域多样化方法在VoxLingua107印度语子集的盲评估中优于传统增强方法。
- 领域泛化显著提升了跨语料库泛化能力,与基线ECAPA-TDNN相比,EER最高降低5.23%。
- 领域不变学习在跨语料库泛化中更为有效,而领域感知学习在同语料库测试集中表现更优。
- 通过多样化增强方法生成的伪领域可实现有效的领域泛化,且无需依赖标注的目标领域数据。
- 语音增强与编解码器类增强等音频增强方法在KGP与LDC测试集上特别有效,显著提升了鲁棒性。
- 尽管性能有所提升,同语料库与跨语料库评估之间仍存在显著性能差距,表明未来通过更真实的增强策略仍有进一步提升空间。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。