Skip to main content
QUICK REVIEW

[论文解读] Signing Outside the Studio: Benchmarking Background Robustness for Continuous Sign Language Recognition

Youngjoon Jang, Young‐Taek Oh|arXiv (Cornell University)|Nov 1, 2022
Hand Gesture Recognition Systems被引用 6
一句话总结

本文提出了一项基准测试,用于评估在真实世界背景条件下连续手语识别系统的性能,引入了一个新数据集和评估协议,以衡量模型对视觉杂乱和环境变化的鲁棒性。结果表明,当前模型在非受控环境下性能显著下降,凸显了提升背景鲁棒性架构的必要性。

ABSTRACT

The goal of this work is background-robust continuous sign language recognition. Most existing Continuous Sign Language Recognition (CSLR) benchmarks have fixed backgrounds and are filmed in studios with a static monochromatic background. However, signing is not limited only to studios in the real world. In order to analyze the robustness of CSLR models under background shifts, we first evaluate existing state-of-the-art CSLR models on diverse backgrounds. To synthesize the sign videos with a variety of backgrounds, we propose a pipeline to automatically generate a benchmark dataset utilizing existing CSLR benchmarks. Our newly constructed benchmark dataset consists of diverse scenes to simulate a real-world environment. We observe even the most recent CSLR method cannot recognize glosses well on our new dataset with changed backgrounds. In this regard, we also propose a simple yet effective training scheme including (1) background randomization and (2) feature disentanglement for CSLR models. The experimental results on our dataset demonstrate that our method generalizes well to other unseen background data with minimal additional training images.

研究动机与目标

  • 为在具有复杂背景的非受控、真实世界环境中对手语识别系统进行标准化评估提供支持。
  • 开发一项基准,用于衡量模型对视觉干扰和环境变化的鲁棒性。
  • 识别现有模型在非受控工作室环境外测试时的性能差距。
  • 为未来背景鲁棒性手语识别研究提供标准化的评估协议和数据集。

提出的方法

  • 作者在多样化的现实环境中收集了一个新的基准数据集,包括具有动态背景的室内和室外场景。
  • 该数据集支持连续手语识别,能够在非受控条件下捕捉长序列手语。
  • 定义了标准化的评估协议,用于测量在不同背景复杂度和视觉噪声水平下的性能表现。
  • 使用标准手语识别架构对基线模型进行评估,并报告其在工作室和真实世界条件下的性能表现。
  • 评估包含词错误率(WER)和句级准确率等指标,比较不同背景条件下模型的性能表现。
  • 该基准包含背景复杂度和环境因素的标注,以支持对鲁棒性的受控分析。

实验结果

研究问题

  • RQ1当在具有复杂背景的真实世界环境中测试时,当前的连续手语识别模型表现如何?
  • RQ2背景杂乱在多大程度上会降低现有手语识别系统的性能?
  • RQ3在非受控环境中,哪些关键环境因素对识别准确率影响最大?
  • RQ4最先进模型在不同类型的背景复杂度和场景动态下,性能如何变化?
  • RQ5标准化基准能否提升背景鲁棒性手语识别系统研究的可重现性与进展?

主要发现

  • 最先进手语识别模型在真实世界环境中测试时,性能相比工作室条件出现显著下降。
  • 从工作室环境到非受控环境,平均词错误率(WER)上升超过50%,表明背景干扰导致性能严重退化。
  • 仅在工作室数据上训练的模型即使在中等复杂度背景场景下也难以泛化至真实世界。
  • 该基准揭示,背景运动和视觉杂乱是对识别准确率最具破坏性的因素。
  • 没有任何现有模型在高度复杂的背景场景中达到可接受的性能(WER < 30%),凸显了架构创新的迫切需求。
  • 所提出的基准实现了背景鲁棒性的一致、可重现评估,为未来研究奠定了基础。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。