東京大学 · 情報科学
Takehiko Ohkawa教授の研究室は、エゴセントリックビジョンや3Dハンドポーズ推定を柱としたコンピュータービジョン分野を専門としています。特に、第一人称視点からの手の3Dポーズ推定や、手と物体のインタラクションを高精度に捉えるためのデータセット構築とアノテーション技術の開発が特徴です。近年では、半教師あり学習やドメイン適応を活用した効率的で高品質なアノテーションパイプラインの構築にも貢献しています。
Figures are computed from collected data and may differ slightly.
We present AssemblyHands, a large-scale benchmark dataset with accurate 3D hand pose annotations, to facilitate the study of egocentric activities with challenging hand-object interactions. The dataset includes synchronized egocentric and exocentric images sampled from the recent Assembly101 dataset, in which participants assemble and disassemble take-apart toys. To obtain high-quality 3D hand pose annotations for the egocentric images, we develop an efficient pipeline, where we use an initial s
Abstract In this survey, we present a systematic review of 3D hand pose estimation from the perspective of efficient annotation and learning. 3D hand pose estimation has been an important research area owing to its potential to enable various applications, such as video understanding, AR/VR, and robotics. However, the performance of models is tied to the quality and quantity of annotated 3D hand poses. Under the status quo, acquiring such annotated 3D hand poses is challenging, e.g., due to the
Hand segmentation is a crucial task in first-person vision. Since first-person images exhibit strong bias in appearance among different environments, adapting a pre-trained segmentation model to a new domain is required in hand segmentation. Here, we focus on appearance gaps for hand regions and backgrounds separately. We propose (i) foreground-aware image stylization and (ii) consensus pseudo-labeling for domain adaptation of hand segmentation. We stylize source images independently for the for
The present study was undertaken to evaluate whether a novel series of 2,6-diaza-5-oxobicyclo[5.4.0]undeca-1(7),8,10-triene derivatives exhibited antagonistic activity for vasopressin V1 and V2 receptors. Most of these compounds were synthesized and showed a high affinity potential for V2 receptor and low to moderate affinity potential for V1 receptor. The most potent and V2-selective compound, N-[4-[2,6-diaza-6-[2-(4-methylpiperazinyl)-2-oxoethyl] -5- oxobicyclo[5.4.0]undeca-1(7),8,10-trien-2-y
Unpaired image-to-image (I2I) translation has received considerable attention in pattern recognition and computer vision because of recent advancements in generative adversarial networks (GANs). However, due to the lack of explicit supervision, unpaired I2I models often fail to generate realistic images, especially in challenging datasets with different backgrounds and poses. Hence, stabilization is indispensable for GANs and applications of I2I translation. Herein, we propose Augmented Cyclic C
We present AssemblyHands, a large-scale benchmark dataset with accurate 3D hand pose annotations, to facilitate the study of egocentric activities with challenging hand-object interactions. The dataset includes synchronized egocentric and exocentric images sampled from the recent Assembly101 dataset, in which participants assemble and disassemble take-apart toys. To obtain high-quality 3D hand pose annotations for the egocentric images, we develop an efficient pipeline, where we use an initial s
Unpaired image-to-image (I2I) translation has received considerable attention in pattern recognition and computer vision because of recent advancements in generative adversarial networks (GANs). However, due to the lack of explicit supervision, unpaired I2I models often fail to generate realistic images, especially in challenging datasets with different backgrounds and poses. Hence, stabilization is indispensable for GANs and applications of I2I translation. Herein, we propose Augmented Cyclic C
Segmentation of lung lobes from MDCT images can provide effective information for functional assessments of each lobe and detection of pulmonary diseases such as the emphysema and lung cancers. Conventional studies have detected pulmonary fissures that located between lung lobes. However, some parts of fissures may disappear in MDCT images because of artifacts or the adhesion between lung lobes. This paper proposes a novel method for segmenting lung lobes based on tubular tissues, which are the
We propose a novel benchmark for cross-view knowledge transfer of dense video captioning, adapting models from web instructional videos with exocentric views to an egocentric view. While dense video captioning (predicting time segments and their captions) is primarily studied with exocentric videos (e.g., YouCook2), benchmarks with egocentric videos are restricted due to data scarcity. To overcome the limited video availability, transferring knowledge from abundant exocentric web videos is deman
In this survey, we present a systematic review of 3D hand pose estimation from the perspective of efficient annotation and learning. 3D hand pose estimation has been an important research area owing to its potential to enable various applications, such as video understanding, AR/VR, and robotics. However, the performance of models is tied to the quality and quantity of annotated 3D hand poses. Under the status quo, acquiring such annotated 3D hand poses is challenging, e.g., due to the difficult
We aim to improve the performance of regressing hand keypoints and segmenting pixel-level hand masks under new imaging conditions (e.g., outdoors) when we only have labeled images taken under very different conditions (e.g., indoors). In the real world, it is important that the model trained for both tasks works under various imaging conditions. However, their variation covered by existing labeled hand datasets is limited. Thus, it is necessary to adapt the model trained on the labeled images (s
Open papers in the app to read, cite, and organize with AI.