Skip to main content
QUICK REVIEW

[論文レビュー] Future Web Growth and its Consequences for Web Search Architectures

Andrew Trotman, Jinglan Zhang|arXiv (Cornell University)|Jul 4, 2013
Web Data Mining and Analysis参考文献 38被引用数 4
ひとこと要約

この論文は、2025年までに1台のハードドライブが全ウェブインデックスを格納可能になると予測し、2035年までに画像を含む検索可能な全ウェブが1台のハードドライブに収容可能になると予測している。本稿では、ユーザーがウェブのローカルスナップショットを保持する分散型の検索アーキテクチャを提案しており、中央集権的なデータセンターへの依存を低減し、オフライン検索を可能にする。代替モデルでは、スター型ネットワークを通じて変更を直接ユーザーにブロードキャストすることで、中央集権的なデータセンターの必要性を排除する。

ABSTRACT

Introduction: Before embarking on the design of any computer system it is first necessary to assess the magnitude of the problem. In the case of a web search engine this assessment amounts to determining the current size of the web, the growth rate of the web, and the quantity of computing resource necessary to search it, and projecting the historical growth of this into the future. Method: The over 20 year history of the web makes it possible to make short-term projections on future growth. The longer history of hard disk drives (and smart phone memory card) makes it possible to make short-term hardware projections. Analysis: Historical data on Internet uptake and hardware growth is extrapolated. Results: It is predicted that within a decade the storage capacity of a single hard drive will exceed the size of the index of the web at that time. Within another decade it will be possible to store the entire searchable text on the same hard drive. Within another decade the entire searchable web (including images) will also fit. Conclusion: This result raises questions about the future architecture of search engines. Several new models are proposed. In one model the user's computer is an active part of the distributed search architecture. They search a pre-loaded snapshot (back-file) of the web on their local device which frees up the online data centre for searching just the difference between the snapshot and the current time. Advantageously this also makes it possible to search when the user is disconnected from the Internet. In another model all changes to all files are broadcast to all users (forming a star-like network) and no data centre is needed.

研究の動機と目的

  • ウェブの長期的成長とその検索エンジンインfra構造に与える影響を評価すること。
  • ウェブおよびハードウェア成長の歴史的傾向に基づいて、将来のストレージおよび計算要件を予測すること。
  • ウェブのストレージが一般消費者用のドライブに収まるようになるに伴い、中央集権的なデータセンターへの依存を低減するためのアーキテクチャ的代替案を検討すること。
  • ローカルデバイスストレージを活用したオフライン検索および分散インデックス化を可能にする検索モデルを設計すること。
  • ピアツーピアによるウェブ変更の配布によって中央集権的なデータセンターを排除することが可能かどうかを評価すること。

提案手法

  • 20年以上にわたるウェブ成長データを外挿し、将来のストレージニーズを予測する。
  • ハードディスクおよびスマートフォンメモリ容量の歴史的傾向を用いて、将来のストレージスケーラビリティを予測する。
  • ユーザーが事前にロードされたローカルスナップショットとしてウェブを保持する分散型検索モデルを提案する。
  • すべてのファイル変更が直接ユーザーにブロードキャストされるスター型ネットワークモデルを導入し、中央集権的なデータセンターの必要性を排除する。
  • オンラインデータセンターが、スナップショットと現在のウェブ状態の差異のみをインデックス化するシステムを設計する。
  • ユーザーがインターネット接続に依存せずにローカルウェブスナップショットを独立して照会できるようにする。

実験結果

リサーチクエスチョン

  • RQ1ウェブの指数的成長は、中央集権的な検索エンジンアーキテクチャのスケーラビリティにどのように影響するか?
  • RQ21台のハードドライブのストレージ容量が、全ウェブインデックスのサイズを超えるのはいつごろか?
  • RQ3ウェブのストレージが一般消費者用ドライブに収まるようになるに伴い、中央集権的なデータセンターへの依存を低減するためのアーキテクチャ的モデルは何か?
  • RQ4ローカルデバイスストレージによるウェブスナップショットの保持によって、オフライン検索を効果的にサポートできるか?
  • RQ5すべてのウェブ変更を直接ユーザーにブロードキャストすることで、中央集権的なデータセンターを排除することが可能か?

主な発見

  • 2025年までに、1台のハードドライブのストレージ容量が全ウェブインデックスのサイズを超えると予測されている。
  • 2035年までに、全ウェブの検索可能なテキストが1台のハードドライブに収容可能になると予測されている。
  • 2045年までに、画像を含む全検索可能なウェブが1台のストレージデバイスに収容可能になると予測されている。
  • ローカルウェブスナップショットを活用する分散型モデルは、オフライン検索を可能にし、中央データセンターの負荷を低減する。
  • ピアツーピアブロードキャストモデルにより、変更を直接ユーザーに配布することで、中央集権的なデータセンターの必要性を排除できる。
  • ローカルストレージおよび分散インデックス化への移行は、伝統的な検索エンジンデータセンターの役割を根本的に変える。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。