Seung-won Hwang
Seoul National University · 情報科学
研究室紹介
Professor Seung-won Hwang's research lab specializes in scalable and efficient data management systems, with a focus on optimizing query processing in large-scale, distributed environments such as web middleware and search engines. The lab investigates ranked and exploratory query processing, information retrieval, and the integration of language models as implicit knowledge bases to address challenges like information overload and high-latency response times. A key research direction involves cost-optimized access strategies for heterogeneous data sources and dynamic result categorization to improve user experience in ad-hoc querying scenarios. The lab also explores the practical deployment and continual updating of large language models for knowledge-intensive applications.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
15Web search engines are optimized to reduce the high-percentile response time to consistently provide fast responses to almost all user queries. This is a challenging task because the query workload exhibits large variability, consisting of many short-running queries and a few long-running queries that significantly impact the high-percentile response time. With modern multicore servers, parallelizing the processing of an individual query is a promising solution to reduce query execution time, bu
Exploratory ad-hoc queries could return too many answers - a phenomenon commonly referred to as "information overload". In this paper, we propose to automatically categorize the results of SQL queries to address this problem. We dynamically generate a labeled, hierarchical category structure - users can determine whether a category is relevant or not by examining simply its label; she can then explore just the relevant categories and ignore the remaining ones, thereby reducing information overlo
We study the problem of supporting ranked queries in middleware environments, where queries are evaluated over multiple sources. In particular, we study Web middleware scenarios, querying over various Web sources. To motivate, consider a Web "travel agent" scenario for finding restaurants and hotels. (We use this real scenario as "benchmark" queries for experiments as well). In particular, how to access sources with different capabilities and costs, to answer queries efficiently? As our Web midd
Language models (LMs) have shown great potential as implicit knowledge bases (KBs). And for their practical use, knowledge in LMs need to be updated periodically. However, existing tasks to assess LMs' efficacy as KBs do not adequately consider multiple large-scale updates.
The explosion of internet usage has provided users with access to information in an unprecedented scale-- The data retrieval problem of finding relevant data has thus become a clear challenge. Such retrieval, with the large scale of data, has naturally demanded ranked answers, or ``best first,'' to enable users to focus on a few top results. \n \nThis thesis presents techniques to support this ranked data retrieval efficiently and effectively First, efficient processing: As data retrieva
This paper studies the problem of open-domain question answering, with the aim of answering a diverse range of questions leveraging knowledge resources. Two types of sources, QA-pair and document corpora, have been actively leveraged with the following complementary strength. The former is highly precise when the paraphrase of given question q was seen and answered during training, often posed as a retrieval problem, while the latter generalizes better for unseen questions. A natural follow-up i
As more and more data are becoming accessible, a naive retrieval of such data may often result in too many answers, as we commonly call "information overload".
Despite the success of neural machine translation models, tensions between fluency of optimizing target language modeling and source-faithfulness remain as challenges. Previously, Conditional Bilingual Mutual Information (CBMI), a scoring metric for the importance of target sentences and tokens, was proposed to encourage fluent and faithful translations. The score is obtained by combining the probability from the translation model and the target language model, which is then used to assign diffe
Research Areas
Seung-won Hwangの研究をNubintでさらに深く
この研究室の論文をアプリで開き、AIと共に読み、要約し、引用しましょう。