Woo-Sung Han
Pohang University of Science and Technology · Computer Science
About the Lab
Professor Woo-Sung Han's research lab specializes in natural language processing, information retrieval, and knowledge graph analytics, with a strong focus on biomedical text mining, entity recognition, and semantic search. The lab develops advanced techniques for extracting and reasoning over complex relationships in scientific literature, including protein-protein interactions and table-text retrieval. It also explores learning-based query optimization and graph processing for large-scale data analytics, aiming to improve accuracy and efficiency in real-world applications.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
8BACKGROUND: Bio-entity extraction is a pivotal component for information extraction from biomedical literature. The dictionary-based bio-entity extraction is the first generation of Named Entity Recognition (NER) techniques. METHODS: This paper presents a hybrid dictionary-based bio-entity extraction technique. The approach expands the bio-entity dictionary by combining different data sources and improves the recall rate through the shortest path edit distance algorithm. In addition, the propose
BACKGROUND: Protein-protein interaction (PPI) extraction has been a focal point of many biomedical research and database curation tools. Both Active Learning and Semi-supervised SVMs have recently been applied to extract PPI automatically. In this paper, we explore combining the AL with the SSL to improve the performance of the PPI task. METHODS: We propose a novel PPI extraction technique called PPISpotter by combining Deterministic Annealing-based SSL and an AL technique to extract protein-pro
Late-interaction based multi-vector retrieval systems have greatly advanced the field of information retrieval by enabling fast and accurate search over millions of documents.However, these systems rely on a naive summation of token-level similarity scores, which often leads to inaccurate relevance estimation caused by the tokenization of semantic units (e.g., words and phrases) and the influence of low-content words (e.g., articles and prepositions).To address these challenges, we propose TRIAL
그래프는 기본적인 데이터 구조 중 하나로, 소셜 네트워크, 단백질 상호 작용 네트워크, 웹 그래프 및 뇌 네트워크와 같은 실 세계의 다양한 응용에서 사용된다. 최근 들어, 소셜 네트워크 기반의 마케팅, 통합 지식 검색 및 인간 커넥톰 분석 등과 같이 빅 그래프 데이터에 대한 분석을 필요로 하는 새로운 서비스 및 기술들의 출현으로 인해, 빅 그래프 데이터를 효율적으로 처리하는 연구에 대한 관심이 증가하고 있다. 본 논문에서는 빅 그래프 데이터 처리 기술들에 대해 살펴본다.
데이터베이스 관리 시스템(DBMS)의 질의 최적화기는 사용자 질의에 적합한 실행 계획을 선정하는 핵심 요소이다. 전통적인 방식은 간단한 추측과 고정된 매개변수를 기반으로 비용을 예측하지만, 이는 실제 실행 비용과 크게 다를 수 있다. 이와 같은 문제점을 극복하기 위해 학습 기반의 질의 최적화 방법들이 연구되어 왔다. 이 논문에서는 일반적인 데이터 환경에서 학습 기반의 비용 예측 모델에 적용한 방법을 우선 소개한다. 이후, 실제 상용 질의 최적화기에 구현하여 TPC-H 벤치마크 하에서의 질의 실행시간을 더욱 줄일 수 있음을 보였고, 추후 개선 사항에 대해 논의한다.
Table-text retrieval aims to retrieve relevant tables and text to support open-domain question answering.Existing studies use either early or late fusion, but face limitations.Early fusion pre-aligns a table row with its associated passages, forming "stars," which often include irrelevant contexts and miss query-dependent relationships.Late fusion retrieves individual nodes, dynamically aligning them, but it risks missing relevant contexts.Both approaches also struggle with advanced reasoning ta
Research Areas
Dive deeper into Woo-Sung Han's research on Nubint
Open this lab's papers in the app to read with AI, summarize, and cite in your writing.