이동원 교수
Dongwon Lee
서울대학교 국사학과 · 컴퓨터과학
연구실 소개
이동원 교수의 연구실은 XML 기반의 데이터 구조 및 스키마 분석, 특히 대규모 디지털 라이브러리에서 발생하는 인용 오류 문제 해결에 초점을 맞추고 있습니다. 정보의 정확성과 품질을 높이기 위한 스케일러블한 알고리즘 기반의 데이터 정제 기법과, e-Commerce 및 SNS 내 사용자 행동 예측 모델링을 통한 디지털 소비 행동 분석도 핵심 연구 분야입니다. 특히 데이터의 구조화, 의미적 제약 조건 통합, 사용자 동기 요인 분석을 기반으로 한 정교한 모델링 기법을 개발하고 있습니다.
연구 현황
연구 성과 추이
표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.
주요 논문
15As XML [5] is emerging as the data format of the internet era, there is an substantial increase of the amount of data in XML format. To better describe such XML data structures and constraints, several XML schema languages have been proposed. In this paper, we present a comparative analysis of six noteworthy XML schema languages.
In this paper, we consider two important problems that commonly occur in bibliographic digital libraries, which seriously degrade their data qualities: Mixed Citation (MC) problem (i.e., citations of different scholars with their names being homonyms are mixed together) and Split Citation (SC) problem (i.e., citations of the same author appear under different name variants). In particular, we investigate an effective yet scalable solution since citations in such digital libraries tend to be larg
Two algorithms, called NeT and CoT, to translate relational schemas to XML schemas using various semantic constraints are presented. The XML schema representation we use is a language-independent formalism named XSchema, that is both precise and concise. A given XSchema can be mapped to a schema in any of the existing XML schema language proposals. Our proposed algorithms have the following characteristics: (1) NeT derives a nested structure from a flat relational model by repeatedly applying th
If they are, only one can refer to a distinct document; if not, many can refer to the same document.
A semantic caching scheme suitable for wrappers wrapping web sources is presented. Since the web sources have typically weaker querying capabilities than conventional databases, existing semantic caching schemes cannot be applied directly. A seamlessly integrated query translation and capability mapping between the wrappers and web sources in semantic caching is described. In addition, an analysis on the match types between the user's input query and cached queries is presented. Semantic knowled
this paper, we study the problems in this conversion. Especially, we are interested in finding XML schema (e.g., DTD, XML-Schema, RELAX) that best describes the existing relational schema. Having the XML schema that precisely describes the semantics and structures of the original relational data is important to further maintain the converted XML documents in future. We first present a straightforward relational to XML translation algorithm, called Flat Translation (FT). Since FT maps the flat re
This dissertation addresses mainly three issues needed to support query relaxation for XML model: framework formalization, extension of existing database techniques, and data conversion between XML and relational models.
The metallic tantalum powder was successfully synthesized via reduction of tantalum pentoxide (Ta2O5) with magnesium gas at 1073~1223 K for 10 h inside the chamber held under an argon atmosphere. The powder obtained after reduction shows the Ta–MgO mixed structure and that the MgO component was dissolved and removed fully via stirring in a water-based HCl solution. The particle size in the tantalum powder obtained after acid leaching was shown to be in a range of 50~300 nm, and the mean internal
In this chapter, three semantics-based schema conversion methods are presented: 1) CPI converts an XML schema to a relational schema while preserving semantic constraints of the original XML schema, 2) NeT derives a nested structured XML schema from a flat relational schema by repeatedly applying the nest operator so that the resulting XML schema becomes hierarchical, and 3) CoT takes a relational schema as input, where multiple tables are interconnected through inclusion dependencies and genera
In this paper, we explore the design space of the Turbo decoding algorithm on GPUs and find a performance bottleneck. We consider three axes for the design space exploration: a radix degree, a parallelization method, and the number of sub-frames per thread block. In Turbo decoding, a degree of radix affects computational complexity and memory access patterns in both algorithmic and implementation viewpoints. Second, computations of branch metrics (BMs) and state metrics (SMs) have a different de
대표 연구 분야
이동원 교수의 연구를 Nubint에서 더 깊이 살펴보세요
이 연구실의 논문을 앱에서 열어 AI와 함께 읽고, 핵심을 요약하고, 내 글에 인용하세요.