Eun‐Sol Kim
Hanyang University · 情報科学
研究室紹介
Professor Eun-Sol Kim's research lab specializes in multimodal representation learning, human-object interaction understanding, and bio-inspired computing systems. The lab develops advanced deep learning frameworks—such as transformer-based architectures and hypergraph attention networks—for structured visual reasoning and cross-modal understanding, with applications in image retrieval, HOI detection, and flexible neuromorphic hardware. A key focus is on modeling complex relationships in visual and biological systems through attention mechanisms, symbolic reasoning, and bio-realistic synaptic plasticity in organic memristors. The lab also explores the integration of biological principles into artificial intelligence and computing systems, bridging neuroscience, computer vision, and materials science.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
15Human-Object Interaction (HOI) detection is a task of identifying "a set of interactions" in an image, which involves the i) localization of the subject (i.e., humans) and target (i.e., objects) of interaction, and ii) the classification of the interaction labels. Most existing methods have indirectly addressed this task by detecting human and object instances and individually inferring every pair of the detected instances. In this paper, we present a novel framework, referred by HOTR, which dir
One of the fundamental problems that arise in multimodal learning tasks is the disparity of information levels between different modalities. To resolve this problem, we propose Hypergraph Attention Networks (HANs), which define a common semantic space among the modalities with symbolic graphs and extract a joint representation of the modalities based on a co-attention map constructed in the semantic space. HANs follow the process: constructing the common semantic space with symbolic graphs of ea
Human-Object Interaction (HOI) detection is the task of identifying a set of (human, object, interaction) triplets from an image. Recent work proposed transformer encoder-decoder architectures that successfully eliminated the need for many hand-designed components in HOI detection through end-to-end training. However, they are limited to single-scale feature resolution, providing suboptimal performance in scenes containing humans, objects, and their interactions with vastly different scales and
Hardware neural networks with mechanical flexibility are promising next-generation computing systems for smart wearable electronics. Several studies have been conducted on flexible neural networks for practical applications; however, developing systems with complete synaptic plasticity for combinatorial optimization remains challenging. In this study, the metal-ion injection density is explored as a diffusive parameter of the conductive filament in organic memristors. Additionally, a flexible ar
Plant growth depends on stem cell niches in meristems. In the root apical meristem, the quiescent center (QC) cells form a niche together with the surrounding stem cells. Stem cells produce daughter cells that are displaced into a transit-amplifying (TA) domain of the root meristem. TA cells divide several times to provide cells for growth. SHORTROOT (SHR) and SCARECROW (SCR) are key regulators of the stem cell niche. Cytokinin controls TA cell activities in a dose-dependent manner. Although the
As a scene graph compactly summarizes the high-level content of an image in a structured and symbolic manner, the similarity between scene graphs of two images reflects the relevance of their contents. Based on this idea, we propose a novel approach for image-to-image retrieval using scene graph similarity measured by graph neural networks. In our approach, graph neural networks are trained to predict the proxy image relevance measure, computed from human-annotated captions using a pre-trained s
Knowledge-based visual question answering (QA) aims to answer a question which requires visually-grounded external knowledge beyond image content itself. Answering complex questions that require multi-hop reasoning under weak supervision is considered as a challenging problem since i) no supervision is given to the reasoning process and ii) highorder semantics of multi-hop knowledge facts need to be captured. In this paper, we introduce a concept of hypergraph to encode highlevel semantics of a
OBJECTIVE: This study aimed to quantitatively confirm the effects of dental specialists' work and stretching on musculoskeletal pain. METHODS: The pain pressure threshold was divided into five parts (neck, shoulder, trunk, lower back, and hand/arm) of the upper body and measured at 15 muscle trigger points. The pain pressure threshold before and after work was measured, and 30 min of stretching and rest were stipulated as an intervention. RESULTS: The pain pressure thresholds reduced significant
Root apical meristem (RAM) drives post-embryonic root growth by constantly supplying cells through mitosis. It is composed of stem cells and their derivatives, the transit-amplifying (TA) cells. Stem cell organization and its maintenance in the RAM are well characterized, however, their relationships with TA cells remain unclear. SHORTROOT (SHR) is critical for root development. It patterns cell types and promotes the post-embryonic root growth. Defective root growth in the shr has been ascribed
Cortical analysis becomes increasingly important for brain research and clinical diagnosis. This problem involves a combinatorial search to find the essential modules among a large number of brain regions. Despite several statistical approaches, cortical analysis remains a formidable challenge due to high dimensionality and sparsity of data. Here we describe an evolutionary method for finding significant modules from cortical data. The method uses a hypernetwork which is encoded as a population
This article is a Letter to the Editor and does not include an Abstract.
In this paper, we consider a problem of analyzing human behavioral data to predict the human cognitive states and generate corresponding actions of sever-agent. Specifically, we aim at predicting human cognitive states during meal time and generating relevant dining services for the human. For this study, we collect behavioral data using 2 kinds of wearable devices, which are an eye tracker and a watch type EDA device, during meal time. We focus on the characteristics of the behavioral data, whi