황승원 교수
Seung-won Hwang
서울대학교 · 컴퓨터과학
연구실 소개
황승원 교수의 연구실은 자연어 처리와 데이터베이스 최적화 분야에서 핵심 기술을 개발하고 있습니다. 특히, 리뷰 분석에서의 쿨스타트 문제 해결을 위한 하이브리드 컨텍스트 기반 감성 분류 기법과, 다양한 소스를 통합해 효율적으로 랭크된 쿼리를 처리하는 미들웨어 기반 쿼리 최적화 기법을 주요 연구 주제로 다룹니다. 또한, 사용자나 제품 정보를 효과적으로 통합해 분류 성능을 향상시키는 메타데이터 처리 기법과 번역 기반 도메인 독립적 문장 맥락 활용 기술 등, 실제 시스템 적용에 초점을 맞춘 지능형 정보 처리 기술을 연구합니다.
연구 현황
연구 성과 추이
표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.
주요 논문
15Web search engines are optimized to reduce the high-percentile response time to consistently provide fast responses to almost all user queries. This is a challenging task because the query workload exhibits large variability, consisting of many short-running queries and a few long-running queries that significantly impact the high-percentile response time. With modern multicore servers, parallelizing the processing of an individual query is a promising solution to reduce query execution time, bu
Exploratory ad-hoc queries could return too many answers - a phenomenon commonly referred to as "information overload". In this paper, we propose to automatically categorize the results of SQL queries to address this problem. We dynamically generate a labeled, hierarchical category structure - users can determine whether a category is relevant or not by examining simply its label; she can then explore just the relevant categories and ignore the remaining ones, thereby reducing information overlo
Software developers increasingly rely on information from the Web, such as documents or code examples on application programming interfaces (APIs), to facilitate their development processes. However, API documents often do not include enough information for developers to fully understand how to use the APIs, and searching for good code examples requires considerable effort. To address this problem, we propose a novel code example recommendation system that combines the strength of browsing docum
We study the problem of supporting ranked queries in middleware environments, where queries are evaluated over multiple sources. In particular, we study Web middleware scenarios, querying over various Web sources. To motivate, consider a Web "travel agent" scenario for finding restaurants and hotels. (We use this real scenario as "benchmark" queries for experiments as well). In particular, how to access sources with different capabilities and costs, to answer queries efficiently? As our Web midd
Language models (LMs) have shown great potential as implicit knowledge bases (KBs). And for their practical use, knowledge in LMs need to be updated periodically. However, existing tasks to assess LMs' efficacy as KBs do not adequately consider multiple large-scale updates.
The explosion of internet usage has provided users with access to information in an unprecedented scale-- The data retrieval problem of finding relevant data has thus become a clear challenge. Such retrieval, with the large scale of data, has naturally demanded ranked answers, or ``best first,'' to enable users to focus on a few top results. \n \nThis thesis presents techniques to support this ranked data retrieval efficiently and effectively First, efficient processing: As data retrieva
This paper studies the problem of open-domain question answering, with the aim of answering a diverse range of questions leveraging knowledge resources. Two types of sources, QA-pair and document corpora, have been actively leveraged with the following complementary strength. The former is highly precise when the paraphrase of given question q was seen and answered during training, often posed as a retrieval problem, while the latter generalizes better for unseen questions. A natural follow-up i
Despite the success of neural machine translation models, tensions between fluency of optimizing target language modeling and source-faithfulness remain as challenges. Previously, Conditional Bilingual Mutual Information (CBMI), a scoring metric for the importance of target sentences and tokens, was proposed to encourage fluent and faithful translations. The score is obtained by combining the probability from the translation model and the target language model, which is then used to assign diffe
As more and more data are becoming accessible, a naive retrieval of such data may often result in too many answers, as we commonly call "information overload".
As the entry and archival of medical data are being digitized, more and more medical data are becoming accessible. This paper studies how to enable an effective retrieval of medical data by ranked retrieval of only the most relevant highly-ranked data. While ranked retrieval has been actively studied lately, existing works have focused mainly on supporting ranking over numerical or text data. However, many existing medical data contain a large amount of categorical attributes, e.g., gender, race
대표 연구 분야
황승원 교수의 연구를 Nubint에서 더 깊이 살펴보세요
이 연구실의 논문을 앱에서 열어 AI와 함께 읽고, 핵심을 요약하고, 내 글에 인용하세요.