The University of Tokyo · Computer Science
Yifei Huang 교수의 연구실은 주로 에고세트릭 비디오(egocentric video) 기반의 행동 이해와 공동 주의 탐지에 초점을 맞추고 있습니다. 특히, 시각적 주의와 행동 인식 간의 상호보완적 관계를 탐색하고, 눈동자 데이터를 활용한 시간적·공간적 주의 지점 탐지 기술을 개발합니다. 또한 행동 분할, 문장 기반 시간적 지정, 소수 샘플에서의 행동 인식 등 비디오 이해의 핵심 과제에 대해 지능형 모델링 기법을 적용합니다. 연구는 주로 그래프 신경망, 조건부 랜덤 필드, 자기지도 학습 등 강력한 표현 학습 기반의 딥러닝 아키텍처를 기반으로 진행됩니다.
Figures are computed from collected data and may differ slightly.
Temporal relations among multiple action segments play an important role in action segmentation especially when observations are limited (e.g., actions are occluded by other objects or happen outside a field of view). In this paper, we propose a network module called Graph-based Temporal Reasoning Module (GTRM) that can be built on top of existing action segmentation models to learn the relation of multiple action segments in various time spans. We model the relations by using two Graph Convolut
In this work, we address two coupled tasks of gaze prediction and action recognition in egocentric videos by exploring their mutual context: the information from gaze prediction facilitates action recognition and vice versa. Our assumption is that during the procedure of performing a manipulation task, on the one hand, what a person is doing determines where the person is looking at. On the other hand, the gaze location reveals gaze regions which contain important and information about the under
The task of weakly supervised temporal sentence grounding aims at finding the corresponding temporal moments of a language description in the video, given video-language correspondence only at video-level. Most existing works select mismatched video-language pairs as negative samples and train the model to generate better positive proposals that are distinct from the negative ones. However, due to the complex temporal structure of videos, proposals distinct from the negative ones may correspond
Joint attention often happens during social interactions, in which individuals share focus on the same object. This article proposes an egocentric vision-based system (ego-vision system) that aims to discover the objects looked at jointly by a group of persons engaged in interactive activities. The proposed system relies on a collection of wearable eye-tracking cameras that provide an egocentric view of the interaction scenes as well as points-of-gaze measurement of each participant. Technically
Abstract The task of few-shot action recognition aims to recognize novel action classes using only a small number of labeled training samples. How to better describe the action in each video and how to compare the similarity between videos are two of the most critical factors in this task. Directly describing the video globally or by its individual frames cannot well represent the spatiotemporal dependencies within an action. On the other hand, naively matching the global representations of two
This work aims to develop a computer-vision technique for understanding objects jointly attended by a group of people during social interactions. As a key tool to discover such objects of joint attention, we rely on a collection of wearable eye-tracking cameras that provide a first-person video of interaction scenes and points-of-gaze data of interacting parties. Technically, we propose a hierarchical conditional random field-based model that can 1) localize events of joint attention temporally
In this work, we address two coupled tasks of gaze prediction and action recognition in egocentric videos by exploring their mutual context. Our assumption is that in the procedure of performing a manipulation task, what a person is doing determines where the person is looking at, and the gaze point reveals gaze and non-gaze regions which contain important and complementary information about the undergoing action. We propose a novel mutual context network (MCN) that jointly learns action-depende
We present a new computational model for gaze prediction in egocentric videos by exploring patterns in temporal shift of gaze fixations (attention transition) that are dependent on egocentric manipulation tasks. Our assumption is that the high-level context of how a task is completed in a certain way has a strong influence on attention transition and should be modeled for gaze prediction in natural dynamic scenes. Specifically, we propose a hybrid model based on deep neural networks which integr
Probiotics play a pivotal role in functional food development owing to their distinct health-promoting properties. This review comprehensively examines probiotics’ classifications and functional mechanisms and their roles in modulating intestinal microbiota, enhancing immunity, and intervening in metabolic diseases. The diverse applications of probiotics in dairy and meat products are examined alongside technological innovations, including microencapsulation, biofilm systems, and personalized st
The human gaze is a cost-efficient physiological data that reveals human underlying attentional patterns. The selective attention mechanism helps the cognition system focus on task-relevant visual clues by ignoring the presence of distractors. Thanks to this ability, human beings can efficiently learn from a very limited number of training samples. Inspired by this mechanism, we aim to leverage gaze for medical image analysis tasks with small training data. Our proposed framework includes a back
In this paper we introduce a provably stable architecture for Neural Ordinary Differential Equations (ODEs) which achieves non-trivial adversarial robustness under white-box adversarial attacks even when the network is trained naturally. For most existing defense methods withstanding strong white-box attacks, to improve robustness of neural networks, they need to be trained adversarially, hence have to strike a trade-off between natural accuracy and adversarial robustness. Inspired by dynamical
In the era of big data, applying big data technology to the community group buying platform can provide a brand new data information processing model for the community group buying supply chain and effectively manage the supply chain costs. This article analyzes and explores the current community group buying supply chain model, and uses data analysis theory to find out the shortcomings of the current model. From the perspective of big data analysis and mining, this article uses big data technol
With today's savvy and empowered customers, sales requires more judgment and becomes more cognitively intense than ever before. We argue that Situation Awareness (SA) is at the center of effective sales and customer engagement in this new era, and Information Fusion (IF) is the key for developing the next generation of decision support systems for digital and AI transformation, leveraging the ubiquitous virtual presence of sales and customer engagement which provides substantially richer capacit
Open papers in the app to read, cite, and organize with AI.