Skip to main content
QUICK REVIEW

[论文解读] Including Signed Languages in Natural Language Processing

Kayo Yin, Amit Moryossef|arXiv (Cornell University)|May 11, 2021
Hand Gesture Recognition Systems参考文献 89被引用 5
一句话总结

本文立场论文主张通过利用语言学洞见、开发专用工具并优先与聋人社群合作,将手语整合到自然语言处理(NLP)中。它呼吁实现标准化分词、基于语言学的模型、大规模真实世界数据收集,并积极吸纳聋人社群参与,以打造公平且高效的手语技术。

ABSTRACT

Signed languages are the primary means of communication for many deaf and hard of hearing individuals. Since signed languages exhibit all the fundamental linguistic properties of natural language, we believe that tools and theories of Natural Language Processing (NLP) are crucial towards its modeling. However, existing research in Sign Language Processing (SLP) seldom attempt to explore and leverage the linguistic organization of signed languages. This position paper calls on the NLP community to include signed languages as a research area with high social and scientific impact. We first discuss the linguistic properties of signed languages to consider during their modeling. Then, we review the limitations of current SLP models and identify the open challenges to extend NLP to signed languages. Finally, we urge (1) the adoption of an efficient tokenization method; (2) the development of linguistically-informed models; (3) the collection of real-world signed language data; (4) the inclusion of local signed language communities as an active and leading voice in the direction of research.

研究动机与目标

  • 解决手语虽为完整自然语言,却在主流NLP中被排除在外的问题。
  • 识别当前手语处理(SLP)模型在利用手语语言结构方面的局限性。
  • 推动开发适配手语视觉-手势性、同步性及空间特性的NLP工具。
  • 通过将聋人社群的需求与领导力置于研发各阶段的核心,确保研究真正惠及聋人社群。
  • 通过可扩展的、社群驱动的数据收集与自动化手段,克服数据稀缺与标注瓶颈。

提出的方法

  • 提出一种针对手语的标准化分词方法,以最小化建模过程中的信息损失。
  • 开发基于语言学洞察的NLP模型,充分考虑手语的句法、形态及空间特性。
  • 从母语手语者处收集大规模真实世界手语数据,避免基于语音的解读方式。
  • 利用可检测帧边界、提取发音特征并支持众包标注的工具,自动化标注流程。
  • 通过包容性研究实践与社群主导的举措,建立与聋人社群的长期、公平合作关系。
  • 将聋人社群的反馈整合至研究设计、评估与发表流程中,以确保其相关性与伦理完整性。

实验结果

研究问题

  • RQ1NLP模型如何适应手语的视觉-手势性、同步性及空间结构特性?
  • RQ2何种标准化分词方法可在手语表征中保留语言信息,同时避免过度数据损失?
  • RQ3如何高效地收集并标注大规模、高质量的手语数据集,同时保持真实性与多样性?
  • RQ4聋人社群在塑造手语技术的发展方向、设计与评估中应扮演何种角色?
  • RQ5如何协同运用NLP与计算机视觉工具,以语言学保真度建模手语?

主要发现

  • 当前SLP模型往往未能有效利用手语的语言结构,而是依赖视觉处理,缺乏语言学基础。
  • 手语数据收集极为耗费资源,每分钟视频可能需要高达600分钟的标注时间。
  • 现有手语数据集在规模、质量和代表性方面均有限,尤其对母语手语者及非解释性内容而言更为不足。
  • 支持帧检测、特征提取与众包的自动化标注工具可显著加速数据管道的开发。
  • 依赖特殊手套或传感器的技术常被聋人社群拒绝,因其被认为体现听觉中心主义且不切实际。
  • 与聋人社群合作对于确保工具满足真实世界需求、避免延续对手语的误解至关重要。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。