Skip to main content
QUICK REVIEW

[論文レビュー] Don't Skype & Type! Acoustic Eavesdropping in Voice-Over-IP

Alberto Compagno, Mauro Conti|arXiv (Cornell University)|Sep 29, 2016
User Authentication and Security Systems参考文献 20被引用数 11
ひとこと要約

本論文は、Skype を通じて送信されるキーストローク音声をキャプチャ・分析することで、リモートでキーボードの音声を盗聴する新規の攻撃手法「Skype & Type (S&T) 攻撃」を紹介する。最小限のプロファイリングで実行可能な本攻撃は、物理的接近や大量のトレーニングデータを必要とせず、実世界の状況でも実現可能である。トップ5の正答率は91.7%に達する。

ABSTRACT

Acoustic emanations of computer keyboards represent a serious privacy issue. As demonstrated in prior work, physical properties of keystroke sounds might reveal what a user is typing. However, previous attacks assumed relatively strong adversary models that are not very practical in many real-world settings. Such strong models assume: (i) adversary's physical proximity to the victim, (ii) precise profiling of the victim's typing style and keyboard, and/or (iii) significant amount of victim's typed information (and its corresponding sounds) available to the adversary. This paper presents and explores a new keyboard acoustic eavesdropping attack that involves Voice-over-IP (VoIP), called Skype & Type (S&T), while avoiding prior strong adversary assumptions. This work is motivated by the simple observation that people often engage in secondary activities (including typing) while participating in VoIP calls. As expected, VoIP software acquires and faithfully transmits all sounds, including emanations of pressed keystrokes, which can include passwords and other sensitive information. We show that one very popular VoIP software (Skype) conveys enough audio information to reconstruct the victim's input -- keystrokes typed on the remote keyboard. Our results demonstrate that, given some knowledge on the victim's typing style and keyboard model, the attacker attains top-5 accuracy of 91.7% in guessing a random key pressed by the victim. Furthermore, we demonstrate that S&T is robust to various VoIP issues (e.g., Internet bandwidth fluctuations and presence of voice over keystrokes), thus confirming feasibility of this attack. Finally, it applies to other popular VoIP software, such as Google Hangouts.

研究の動機と目的

  • Skype などの VoIP ソフトウェアを介したリモートキーボード音声盗聴の可能性を調査すること。キーストローク音声が通話中に意図せず送信される点を対象とする。
  • 被害者のタイピングスタイルやキーボードの詳細なプロファイリングを必要とする、従来の攻撃手法の限界を是正すること。
  • 帯域幅の変動やキーストロークに重なった音声干渉を含む、実世界の VoIP 条件下での攻撃の耐性を評価すること。
  • スペクトル特徴に基づくキーストローク推定攻撃に対する対策を検討すること。

提案手法

  • 本攻撃は、Skype などの VoIP ソフトウェアが通話中にすべての音声(キーストローク音声を含む)をキャプチャ・送信することを利用している。
  • 捕獲した音声ストリームを分析するために、メル周波数ケプストラム係数(MFCC)と高速フーリエ変換(FFT)特徴を組み合わせて使用する。
  • 被害者のキーボードと同じモデルのキーボードからの限られたキーストローク音声データを用いて機械学習分類器をトレーニングし、被害者固有のプロファイリングの必要を最小限に抑える。
  • 音声圧縮、ダウンサンプリング、モノラルチャンネル混合などの VoIP 特有の歪みに対する耐性を評価する。
  • 分類に使用されるスペクトル特徴を攪乱するために、ランダムなマルチバンドイコライゼーションを対策としてテストする。
  • 複数のラップトップとユーザーを対象に10分割交差検証を実施し、正答率と一般化性能を検証する。

実験結果

リサーチクエスチョン

  • RQ1被害者のデバイスに物理的アクセスがなくても、VoIP 音声ストリームからキーストローク音声を信頼性高く抽出・分類できるか?
  • RQ2同じキーボードモデルのプロファイリングのみで、限られたトレーニングデータでのキーストローク推定の正確性はどの程度か?
  • RQ3音声圧縮、ダウンサンプリング、モノラル混合などの VoIP 特有の信号処理は、音声盗聴の実現可能性にどのように影響するか?
  • RQ4ランダムなイコライゼーションやその他の対策によって、スペクトル特徴に基づく攻撃はどの程度軽減可能か?
  • RQ5S&T 攻撃は Skype 以外の VoIP プラットフォーム(例:Google Hangouts)に対しても適用可能か?

主な発見

  • S&T 攻撃は、被害者が押したキーボードのランダムなキーを推定する際、キーボードモデルのプロファイリングと最小限のトレーニングデータでのみ、トップ5正答率91.7%を達成する。
  • 帯域幅の制限やキーストローク中に人の声が重なっている状況でも、攻撃は耐性を示す。
  • ランダムなマルチバンドイコライゼーションは FFT に基づく特徴を顕著に攪乱し、攻撃の正答率をランダム推測レベルまで低下させる。
  • MFCC に基づく特徴はイコライゼーションに対し部分的に耐性を示すため、このような対策後でもスペクトル特徴に基づく攻撃は依然として実現可能である。
  • 初期の実験では、S&T 攻撃が Google Hangouts に対しても同様に実現可能であることが確認され、他の VoIP プラットフォームへの広範な適用可能性が示唆される。
  • 本研究は、最小限の攻撃者能力でも、VoIP を介したリモート音声盗聴が現実の脅威であることを示している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。