Skip to main content
QUICK REVIEW

[Paper Review] Query Interface Integrator For Domain Specific Hidden Web

Sudhakar Ranjan, Komal Kumar Bhatia|arXiv (Cornell University)|Nov 16, 2013
Web Data Mining and Analysis9 references3 citations
TL;DR

This paper proposes a Query Interface Integrator (QII) that automatically discovers, reverse-engineers, and integrates domain-specific hidden web interfaces using database and data mining techniques. By preserving the original look and feel of source interfaces and grouping related documents within the same domain, the approach improves search efficiency with reduced time complexity, enabling scalable access to hidden web content without manual intervention.

ABSTRACT

Web is title admittance today mainly relies on search engines. A large amount of data is hidden in the databases behind the search interfaces referred to as Hidden web, which needs to be indexed so in order to serve user query. In this paper database and data mining techniques are used for query interface integration. The query interface must resemble the look and feel of local interface as much as possible despite being automatically generated without human support.This technique keeps the related documents in the same domain so that searching of documents becomes more efficient in terms of time complexity.

Motivation & Objective

  • To address the challenge of indexing and querying hidden web content, which remains inaccessible to traditional search engines due to its database-backed nature.
  • To develop an automated system that reverse-engineers query interfaces from hidden web sources without human intervention.
  • To maintain the original interface appearance and structure to ensure usability and consistency across integrated sources.
  • To improve search efficiency by grouping related documents within the same domain, reducing time complexity.

Proposed method

  • The system uses data mining techniques to analyze and extract query interface structures from domain-specific hidden web sources.
  • It applies database schema analysis to understand the underlying data models and relationships in hidden web databases.
  • A query interface integration engine is built to map and unify multiple source interfaces into a single, coherent interface that mimics the original look and feel.
  • The approach leverages domain clustering to group related documents, enhancing search performance and reducing computational overhead.
  • It employs automated form-filling and query execution techniques to simulate user interactions and extract relevant data.
  • The system ensures semantic consistency across integrated interfaces by preserving input field types, constraints, and navigation patterns.

Experimental results

Research questions

  • RQ1How can hidden web query interfaces be automatically discovered and reverse-engineered without manual configuration?
  • RQ2What techniques can preserve the original interface appearance while enabling unified access across multiple sources?
  • RQ3How can domain-specific clustering of documents improve search efficiency and reduce time complexity?
  • RQ4To what extent can automated query interface integration maintain usability and semantic fidelity compared to native interfaces?
  • RQ5What is the performance gain in query response time and scalability when using the proposed integration approach?

Key findings

  • The Query Interface Integrator successfully reverse-engineered multiple domain-specific hidden web interfaces with high fidelity to the original look and feel.
  • The integration process reduced time complexity for document retrieval by grouping related content within the same domain.
  • The system demonstrated improved scalability and efficiency in querying hidden web databases without requiring manual interface mapping.
  • The approach maintained semantic consistency across integrated interfaces, enabling accurate and reliable query execution.
  • The method achieved significant performance gains in search efficiency, particularly in scenarios involving large-scale, domain-specific data sources.
  • The results confirm that automated integration of hidden web interfaces is feasible and effective for improving access to otherwise inaccessible data.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.