Skip to main content

Hanbyul Joo

Seoul National University · 情報科学

研究室紹介

Professor Hanbyul Joo's research lab specializes in 3D human motion and shape estimation, focusing on developing advanced deep learning methods to reconstruct detailed 3D human bodies and their motions from monocular images or video. The lab addresses key challenges such as occlusion, large-scale social interactions, and high-fidelity modeling of full-body motion—including face and fingers—using multi-view systems and novel neural representations. A central theme is the creation of large-scale, ethical, and diverse datasets like Ego4D to enable robust and generalizable models for egocentric and social interaction scenarios.

3D human motionmonocular reconstructionegocentric videosocial interaction3D human shape

Research Overview

Papers
85
Total Citations
3,766
Papers (5y)
37
Primary Field
情報科学

Research Output Trend

Figures are computed from collected data and may differ slightly.

Publications per year (5y)
37total
2022
2023
2024
2025
2026
Citations per year (5y)
809total
20222023202420252026

Selected Papers

15
1
Article|762 citations·2020
PIFuHD: Multi-Level Pixel-Aligned Implicit Function for High-Resolution 3D Human Digitization
Shunsuke Saito, Tomas Simon, Jason Saragih, Hanbyul Joo

Recent advances in image-based 3D human shape estimation have been driven by the significant improvement in representation power afforded by deep neural networks. Although current approaches have demonstrated the potential in real world settings, they still fail to produce reconstructions with the level of detail often present in the input images. We argue that this limitation stems primarily form two conflicting requirements; accurate predictions require large context, but precise predictions r

Computer Vision and Pattern RecognitionComputer Science
2
Article|535 citations·2015
Panoptic Studio: A Massively Multiview System for Social Motion Capture
Hanbyul Joo, Hao Liu, Lei Tan, Lin Gui, Bart Nabbe, Iain Matthews, Takeo Kanade, Shohei Nobuhara, Yaser Sheikh

We present an approach to capture the 3D structure and motion of a group of people engaged in a social interaction. The core challenges in capturing social interactions are: (1) occlusion is functional and frequent, (2) subtle motion needs to be measured over a space large enough to host a social group, and (3) human appearance and configuration variation is immense. The Panoptic Studio is a system organized around the thesis that social interactions should be measured through the perceptual int

Computer Vision and Pattern RecognitionComputer Science
3
Article|487 citations·2022
Ego4D: Around the World in 3,000 Hours of Egocentric Video
Kristen Grauman, Andrew Westbury, Eugene H. Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, Miguel Martin, Tushar Nagarajan
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)OA

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of dailylife activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countries. The approach to collection is designed to uphold rigorous privacy and ethics standards, with consenting participants and robust de-identification procedures where relevant. Ego4D dramatically expan

Computer Vision and Pattern RecognitionComputer Science
4
Article|353 citations·2017
Panoptic Studio: A Massively Multiview System for Social Interaction Capture
Hanbyul Joo, Tomas Simon, Xulong Li, Hao Liu, Lei Tan, Lin Gui, Sean Banerjee, Timothy Godisart, Bart Nabbe, Iain Matthews, Takeo Kanade, Shohei Nobuhara
SJR Q1IEEE Transactions on Pattern Analysis and Machine Intelligence

We present an approach to capture the 3D motion of a group of people engaged in a social interaction. The core challenges in capturing social interactions are: (1) occlusion is functional and frequent; (2) subtle motion needs to be measured over a space large enough to host a social group; (3) human appearance and configuration variation is immense; and (4) attaching markers to the body may prime the nature of interactions. The Panoptic Studio is a system organized around the thesis that social

Computer Vision and Pattern RecognitionComputer Science
5
Article|341 citations·2019
Monocular Total Capture: Posing Face, Body, and Hands in the Wild
Donglai Xiang, Hanbyul Joo, Yaser Sheikh

We present the first method to capture the 3D total motion of a target person from a monocular view input. Given an image or a monocular video, our method reconstructs the motion from body, face, and fingers represented by a 3D deformable mesh model. We use an efficient representation called 3D Part Orientation Fields (POFs), to encode the 3D orientations of all body parts in the common 2D image space. POFs are predicted by a Fully Convolutional Network, along with the joint confidence maps. To

Computer Vision and Pattern RecognitionComputer Science
6
Article|181 citations·2021
Exemplar Fine-Tuning for 3D Human Model Fitting Towards In-the-Wild 3D Human Pose Estimation
Hanbyul Joo, Natalia Neverova, Andrea Vedaldi
2021 International Conference on 3D Vision (3DV)

Differently from 2D image datasets such as COCO, largescale human datasets with 3D ground-truth annotations are very difficult to obtain in the wild. In this paper, we address this problem by augmenting existing 2D datasets with high-quality 3D pose fits. Remarkably, the resulting annotations are sufficient to train from scratch 3D pose regressor networks that outperform the current state-of-the-art on in the-wild benchmarks such as 3DPW. Additionally, training on our augmented data is straightf

Computer Vision and Pattern RecognitionComputer Science
7
Preprint|140 citations·2017
Hand Keypoint Detection in Single Images Using Multiview Bootstrapping
Tomas Simon, Hanbyul Joo, Iain Matthews, Yaser Sheikh
OA

We present an approach that uses a multi-camera system to train fine-grained detectors for keypoints that are prone to occlusion, such as the joints of a hand. We call this procedure multiview bootstrapping: first, an initial keypoint detector is used to produce noisy labels in multiple views of the hand. The noisy detections are then triangulated in 3D using multiview geometry or marked as outliers. Finally, the reprojected triangulations are used as new labeled training data to improve the det

Computer Vision and Pattern RecognitionComputer Science
8
Article|108 citations·2022
BANMo: Building Animatable 3D Neural Models from Many Casual Videos
Gengshan Yang, Minh Vo, Natalia Neverova, Deva Ramanan, Andrea Vedaldi, Hanbyul Joo
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

Prior work for articulated 3D shape reconstruction often relies on specialized multi-view and depth sensors or pre-built deformable 3D models. Such methods do not scale to diverse sets of objects in the wild. We present a method that requires neither of them. It aims to create high-fidelity, articulated 3D models from many casual RGB videos in a differentiable rendering framework. Our key in-sight is to merge three schools of thought: (1) classic deformable shape models that make use of articula

Computational MechanicsEngineering
9
Book Chapter|106 citations·2020
Perceiving 3D Human-Object Spatial Arrangements from a Single Image in the Wild
Jason Zhang, Sam Pepose, Hanbyul Joo, Deva Ramanan, Jitendra Malik, Angjoo Kanazawa
SJR Q2Lecture notes in computer science
Computational MechanicsEngineering
10
Article|85 citations·2022
Learning to Listen: Modeling Non-Deterministic Dyadic Facial Motion
Evonne Ng, Hanbyul Joo, Liwen Hu, Hao Li, Trevor Darrell, Angjoo Kanazawa, Shiry Ginosar
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

We present a framework for modeling interactional communication in dyadic conversations: given multimodal inputs of a speaker, we autoregressively output multiple possibilities of corresponding listener motion. We combine the motion and speech audio of the speaker using a motion-audio cross attention transformer. Furthermore, we enable non-deterministic prediction by learning a discrete latent representation of realistic listener motion with a novel motion-encoding VQ-VAE. Our method organically

Signal ProcessingComputer Science
11
Article|69 citations·2019
Single-Network Whole-Body Pose Estimation
Gines Hidalgo Martinez, Yaadhav Raaj, Haroon Idrees, Donglai Xiang, Hanbyul Joo, Tomas Simon, Yaser Sheikh

We present the first single-network approach for 2D~whole-body pose estimation, which entails simultaneous localization of body, face, hands, and feet keypoints. Due to the bottom-up formulation, our method maintains constant real-time performance regardless of the number of people in the image. The network is trained in a single stage using multi-task learning, through an improved architecture which can handle scale differences between body/foot and face/hand keypoints. Our approach considerabl

Computer Vision and Pattern RecognitionComputer Science
12
Preprint|54 citations·2018
Total Capture: A 3D Deformation Model for Tracking Faces, Hands, and Bodies
Hanbyul Joo, Tomas Simon, Yaser Sheikh
OA

The following topics are dealt with: learning (artificial intelligence); feature extraction; image classification; neural nets; image representation; object detection; image segmentation; convolution; feedforward neural nets; video signal processing.

Computer Vision and Pattern RecognitionComputer Science
13
Preprint|51 citations·2020
Exemplar Fine-Tuning for 3D Human Pose Fitting Towards In-the-Wild 3D Human Pose Estimation
Hanbyul Joo, Natalia Neverova, Andrea Vedaldi
arXiv (Cornell University)OA

We propose a method for building large collections of human poses with full 3D annotations captured `in the wild', for which specialized capture equipment cannot be used. We start with a dataset with 2D keypoint annotations such as COCO and MPII and generates corresponding 3D poses. This is done via Exemplar Fine-Tuning (EFT), a new method to fit a 3D parametric model to 2D keypoints. EFT is accurate and can exploit a data-driven pose prior to resolve the depth reconstruction ambiguity that come

Computer Vision and Pattern RecognitionComputer Science
14
Article|35 citations·2014
MAP Visibility Estimation for Large-Scale Dynamic 3D Reconstruction
Hanbyul Joo, Hyun Soo Park, Yaser Sheikh

Many traditional challenges in reconstructing 3D motion, such as matching across wide baselines and handling occlusion, reduce in significance as the number of unique viewpoints increases. However, to obtain this benefit, a new challenge arises: estimating precisely which cameras observe which points at each instant in time. We present a maximum a posteriori (MAP) estimate of the time-varying visibility of the target points to reconstruct the 3D motion of an event from a large number of cameras.

Computer Vision and Pattern RecognitionComputer Science
15
Preprint|16 citations·2016
Panoptic Studio: A Massively Multiview System for Social Interaction Capture
Hanbyul Joo, Tomas Simon, Xulong Li, Hao Liu, Lei Tan, Lin Gui, Sean Banerjee, Timothy Godisart, Bart Nabbe, Iain Matthews, Takeo Kanade, Shohei Nobuhara
arXiv (Cornell University)OA

We present an approach to capture the 3D motion of a group of people engaged\nin a social interaction. The core challenges in capturing social interactions\nare: (1) occlusion is functional and frequent; (2) subtle motion needs to be\nmeasured over a space large enough to host a social group; (3) human appearance\nand configuration variation is immense; and (4) attaching markers to the body\nmay prime the nature of interactions. The Panoptic Studio is a system organized\naround the thesis that s

Computer Vision and Pattern RecognitionComputer Science

Research Areas

Computer Vision and Pattern RecognitionControl and Systems EngineeringComputational MechanicsSignal ProcessingBiomedical EngineeringSociology and Political Science

Hanbyul Jooの研究をNubintでさらに深く

この研究室の論文をアプリで開き、AIと共に読み、要約し、引用しましょう。