[논문 리뷰] Future Web Growth and its Consequences for Web Search Architectures
이 논문은 2025년까지 단일 하드디스크가 전체 웹 인덱스를 저장할 것으로 예측하며, 2035년에는 이미지까지 포함한 전체 검색 가능한 웹도 하나의 디스크에 담길 것으로 보고 있다. 사용자가 웹의 로컬 스냅샷을 보관하는 분산 검색 아키텍처를 제안하여 중앙집중식 데이터 센터에 대한 의존도를 줄이고 오프라인 검색을 가능하게 하며, 다른 모델에서는 성능 네트워크를 통해 사용자에게 직접 변경 사항을 브로드캐스트한다.
Introduction: Before embarking on the design of any computer system it is first necessary to assess the magnitude of the problem. In the case of a web search engine this assessment amounts to determining the current size of the web, the growth rate of the web, and the quantity of computing resource necessary to search it, and projecting the historical growth of this into the future. Method: The over 20 year history of the web makes it possible to make short-term projections on future growth. The longer history of hard disk drives (and smart phone memory card) makes it possible to make short-term hardware projections. Analysis: Historical data on Internet uptake and hardware growth is extrapolated. Results: It is predicted that within a decade the storage capacity of a single hard drive will exceed the size of the index of the web at that time. Within another decade it will be possible to store the entire searchable text on the same hard drive. Within another decade the entire searchable web (including images) will also fit. Conclusion: This result raises questions about the future architecture of search engines. Several new models are proposed. In one model the user's computer is an active part of the distributed search architecture. They search a pre-loaded snapshot (back-file) of the web on their local device which frees up the online data centre for searching just the difference between the snapshot and the current time. Advantageously this also makes it possible to search when the user is disconnected from the Internet. In another model all changes to all files are broadcast to all users (forming a star-like network) and no data centre is needed.
연구 동기 및 목표
- 웹의 장기적 성장을 및 검색 엔진 인프라에 대한 영향을 평가하기 위해.
- 과거 웹 및 하드웨어 성장 추세를 바탕으로 향후 스토리지 및 계산 요구사항을 예측하기 위해.
- 웹 스토리지가 소비자용 드라이브에 맞을 때 중앙집중식 데이터 센터에 대한 의존도를 줄일 수 있는 아키텍처 모델을 탐색하기 위해.
- 로컬 장치 스토리지 기반으로 오프라인 검색과 분산 색인화를 가능하게 하는 검색 모델을 설계하기 위해.
- 피어 투 피어 방식으로 웹 변경 사항을 배포함으로써 중앙집중식 데이터 센터를 제거할 수 있는지 타당성을 평가하기 위해.
제안 방법
- 20년 이상의 웹 성장 데이터를 외삽하여 향후 스토리지 수요를 예측하기 위해.
- 하드디스크 및 스마트폰 메모리 용량의 역사적 추세를 활용해 향후 스토리지 확장성 예측하기 위해.
- 사용자가 사전 로드된 웹의 로컬 스냅샷을 유지하는 분산 검색 모델 제안하기 위해.
- 모든 파일 변경 사항이 직접 사용자에게 브로드캐스트되는 성능 네트워크 유사 모델을 도입하여 중앙집중식 데이터 센터가 필요 없도록 하기 위해.
- 온라인 데이터 센터가 스냅샷과 현재 웹 상태의 차이만 색인화하도록 시스템 설계하기 위해.
- 인터넷 연결 없이도 사용자가 로컬 웹 스냅샷을 독립적으로 쿼리할 수 있도록 하여 오프라인 검색 지원하기 위해.
실험 결과
연구 질문
- RQ1웹의 기하급수적 성장이 중앙집중식 검색 엔진 아키텍처의 확장성에 어떤 영향을 미칠 것인가?
- RQ2언제부터 단일 하드디스크의 스토리지 용량이 전체 웹 인덱스의 크기를 초과할 것인가?
- RQ3웹 스토리지가 소비자용 드라이브에 맞을 때 중앙집중식 데이터 센터 의존도를 줄일 수 있는 아키텍처 모델은 무엇인가?
- RQ4로컬 장치 스토리지 기반의 웹 스냅샷으로 오프라인 검색을 효과적으로 지원할 수 있는가?
- RQ5사용자에게 직접 웹 변경 사항을 브로드캐스트함으로써 중앙집중식 데이터 센터를 제거하는 것은 타당한가?
주요 결과
- 2025년까지 단일 하드디스크의 스토리지 용량이 전체 웹 인덱스의 크기를 초과할 것으로 예측된다.
- 2035년까지 전체 검색 가능한 웹 텍스트가 하나의 하드디스크에 담길 것으로 예측된다.
- 2045년까지 전체 검색 가능한 웹(이미지 포함)이 하나의 스토리지 장치에 담길 것으로 예측된다.
- 로컬 웹 스냅샷을 활용한 분산 모델은 오프라인 검색을 가능하게 하고 중앙 데이터 센터의 부담을 줄인다.
- 피어 투 피어 브로드캐스트 모델은 사용자에게 직접 변경 사항을 배포함으로써 중앙집중식 데이터 센터가 필요 없도록 할 수 있다.
- 로컬 스토리지 및 분산 색인화로의 전환은 전통적인 검색 엔진 데이터 센터의 역할을 근본적으로 변화시킬 것이다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.