Skip to main content
QUICK REVIEW

[论文解读] DolphinAtack: Inaudible Voice Commands

Guoming Zhang, Chen Yan|arXiv (Cornell University)|Aug 31, 2017
Speech and Audio Processing参考文献 29被引用 126
一句话总结

DolphinAttack 显示,无法听见的超声波语音命令可以通过非线性硬件被 MEMS/ECM 麦克风解调,并被流行的 SR 系统解读,从而在多个设备上实现隐身式激活与控制指令。它还提出了防御方法。

ABSTRACT

Speech recognition (SR) systems such as Siri or Google Now have become an increasingly popular human-computer interaction method, and have turned various systems into voice controllable systems(VCS). Prior work on attacking VCS shows that the hidden voice commands that are incomprehensible to people can control the systems. Hidden voice commands, though hidden, are nonetheless audible. In this work, we design a completely inaudible attack, DolphinAttack, that modulates voice commands on ultrasonic carriers (e.g., f > 20 kHz) to achieve inaudibility. By leveraging the nonlinearity of the microphone circuits, the modulated low frequency audio commands can be successfully demodulated, recovered, and more importantly interpreted by the speech recognition systems. We validate DolphinAttack on popular speech recognition systems, including Siri, Google Now, Samsung S Voice, Huawei HiVoice, Cortana and Alexa. By injecting a sequence of inaudible voice commands, we show a few proof-of-concept attacks, which include activating Siri to initiate a FaceTime call on iPhone, activating Google Now to switch the phone to the airplane mode, and even manipulating the navigation system in an Audi automobile. We propose hardware and software defense solutions. We validate that it is feasible to detect DolphinAttack by classifying the audios using supported vector machine (SVM), and suggest to re-design voice controllable systems to be resilient to inaudible voice command attacks.

研究动机与目标

  • 演示针对语音可控系统(VCS)的无法听见的语音命令注入的可行性。
  • 解释超声调制如何被麦克风非线性解调以到达语音识别系统(SR)的方法。
  • 在多种 SR 系统和设备平台上验证攻击,以评估安全影响。
  • 提出硬件/软件防御措施以减轻无法听见的命令攻击。

提出的方法

  • 使用幅度调制(AM)将基带语音命令调制到超声载波上。
  • 利用麦克风非线性在低通滤波器(LPF)阶段之前解调并恢复基带命令。
  • 表征麦克风非线性,并展示 MEMS 和 ECM 麦克风的解调。
  • 设计实际发射器(桌面式和基于便携智能手机的)以注入无法听见的命令。
  • 在多样的 SR 系统(Siri、Google Now、Alexa 等)上评估激活和通用控制命令。
  • 评估载波频率选择、调制深度和语音选择以优化攻击。

实验结果

研究问题

  • RQ1普通麦克风硬件是否能够解调无法听见的超声命令并被 SR 系统解码?
  • RQ2哪些硬件和软件因素会影响 DolphinAttack 在不同设备上的成功?
  • RQ3在主流 SR 平台上,可以通过无法听见的注入实现哪些激活和控制命令?
  • RQ4哪些防御措施可以有效检测或减轻此类无法听见的命令攻击?

主要发现

  • DolphinAttack 能在超声载波上注入无法听见的命令,并被 SR 系统如 Siri、Google Now、Alexa 解码。
  • 即使没有所有者的语音样本,也可生成激活命令,在测试的 89 种激活命令类型中实现 39% 的成功率(35 次成功)。
  • 攻击在 7 个 SR 系统和 16 个设备/平台上得到验证。
  • 实验表明通过麦克风非线性,MEMS 和 ECM 麦克风都成功解调基带信号。
  • 研究包括实际发射器设计(桌面式和便携式)并展示潜在的现实世界攻击场景(如 FaceTime、飞行模式、导航)。
  • 作者提出硬件/软件防御措施以缓解 DolphinAttack。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。