Skip to main content
QUICK REVIEW

[论文解读] Unacceptable, where is my privacy? Exploring Accidental Triggers of Smart Speakers

Lea Schönherr, Maximilian Golla|arXiv (Cornell University)|Aug 2, 2020
User Authentication and Security Systems参考文献 37被引用 16
一句话总结

本文提出一种自动化方法,利用基于音素的加权Levenshtein距离和TTS生成的音频,系统性地识别来自8家制造商的11款智能扬声器中意外触发词——即被误识别为唤醒词的声音。研究识别出超过1,000个经验证的意外触发词,揭示了因错误唤醒词激活导致的重大隐私风险,并倡导通过提升透明度、引入安全词(safewords)和加强隐私控制措施来减轻意外数据收集问题。

ABSTRACT

Voice assistants like Amazon's Alexa, Google's Assistant, or Apple's Siri, have become the primary (voice) interface in smart speakers that can be found in millions of households. For privacy reasons, these speakers analyze every sound in their environment for their respective wake word like ''Alexa'' or ''Hey Siri,'' before uploading the audio stream to the cloud for further processing. Previous work reported on the inaccurate wake word detection, which can be tricked using similar words or sounds like ''cocaine noodles'' instead of ''OK Google.'' In this paper, we perform a comprehensive analysis of such accidental triggers, i.,e., sounds that should not have triggered the voice assistant, but did. More specifically, we automate the process of finding accidental triggers and measure their prevalence across 11 smart speakers from 8 different manufacturers using everyday media such as TV shows, news, and other kinds of audio datasets. To systematically detect accidental triggers, we describe a method to artificially craft such triggers using a pronouncing dictionary and a weighted, phone-based Levenshtein distance. In total, we have found hundreds of accidental triggers. Moreover, we explore potential gender and language biases and analyze the reproducibility. Finally, we discuss the resulting privacy implications of accidental triggers and explore countermeasures to reduce and limit their impact on users' privacy. To foster additional research on these sounds that mislead machine learning models, we publish a dataset of more than 1000 verified triggers as a research artifact.

研究动机与目标

  • 系统分析并量化跨多种智能扬声器的意外触发词——即被错误分类为唤醒词的声音。
  • 评估现有防护机制(如基于云的验证和声学指纹)在减轻意外触发词方面的作用效果。
  • 开发并验证一种自动化方法,利用音素距离度量来生成和检测潜在的意外触发词。
  • 探讨意外唤醒词激活带来的隐私影响,并提出可操作的应对措施。

提出的方法

  • 作者使用发音词典和一种新型基于音素的Levenshtein距离(采用音素相关权重)来生成潜在的意外触发词候选。
  • 他们利用文本到语音(TTS)服务为这些候选词合成音频,以在真实智能扬声器上测试检测效果。
  • 通过受控实验设置模拟包含电视节目、新闻和专业音频数据集等多种音源的客厅环境。
  • 该方法评估了本地和基于云的唤醒词检测系统,以确定触发词是否能绕过各层验证机制。
  • 研究人员在一致条件下对11款来自8家制造商的智能扬声器进行基准测试,以确保可复现性和公平性。
  • 他们发布了一个包含1,000多个经验证的意外触发词的数据集,以支持未来对语音助手对抗性输入的研究。

实验结果

研究问题

  • RQ1不同智能扬声器中意外触发词的普遍程度如何?哪些因素影响其检测率?
  • RQ2基于云的验证和声学指纹技术在多大程度上降低了意外触发词的影响?
  • RQ3基于音素的加权Levenshtein距离能否有效识别出可能被误分类为唤醒词的声音?
  • RQ4语言、性别和音频内容如何影响意外触发词的发生概率?
  • RQ5有哪些实际可行的对策可以降低因意外唤醒词激活带来的隐私风险?

主要发现

  • 作者通过一种自动化且基于音素的识别方法,在来自8家制造商的11款智能扬声器中成功识别出超过1,000个经验证的意外触发词。
  • 使用基于音素且带权重的Levenshtein距离显著优于标准Levenshtein度量,在识别合理意外触发词方面表现更优。
  • 所有测试的智能扬声器均存在意外触发词,表明这是主流厂商中普遍存在且系统性的问题。
  • 基于云的验证和声学指纹技术虽能减少但无法完全消除误报,说明本地检测仍是薄弱环节。
  • 意外触发词更可能出现在与常见唤醒词具有相似起始或结尾音素的语音结构中。
  • 本研究表明,当前的隐私控制措施(如静音按钮)不足以应对风险,并提出更安全的替代方案,如引入安全词(safewords)和提升录音管理的透明度。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。