Skip to main content

주한별 교수

Hanbyul Joo

서울대학교 컴퓨터공학부 · 컴퓨터과학

연구실 소개

주한별 교수의 연구실은 3D 인간 자세 추정 및 운동 추적 기술에 중점을 두고 있으며, 특히 단일 카메라 영상에서 신체 전체의 정밀한 3D 모션을 복원하는 데 혁신적인 기여를 하고 있습니다. 다양한 시각적 관점과 고해상도 정보를 융합해 복잡한 상호작용, 오염, 다양한 체형 변화에도 강건한 분석을 가능하게 하는 다중 카메라 시스템과 egocentric(자기 중심) 비디오 기반 행동 이해 연구도 함께 진행하고 있습니다. 특히, 실제 환경에서의 다양성과 윤리적 기준을 고려한 대규모 3D 데이터셋 구축과 이를 활용한 신경망 학습 프로토콜 개발이 핵심입니다.

3D 인간 모델링단일 카메라 운동 추정다중 시점 통합에고세트릭 비디오대규모 3D 데이터셋

연구 현황

논문 수
85
총 인용 수
3,766
최근 5년 논문
37
주요 분야
컴퓨터과학

연구 성과 추이

표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.

5개년 연도별 논문 게재 수
37총합
2022
2023
2024
2025
2026
5개년 연도별 피인용 수
809총합
20222023202420252026

주요 논문

15
1
논문|인용수 762·2020
PIFuHD: Multi-Level Pixel-Aligned Implicit Function for High-Resolution 3D Human Digitization
Shunsuke Saito, Tomas Simon, Jason Saragih, Hanbyul Joo

Recent advances in image-based 3D human shape estimation have been driven by the significant improvement in representation power afforded by deep neural networks. Although current approaches have demonstrated the potential in real world settings, they still fail to produce reconstructions with the level of detail often present in the input images. We argue that this limitation stems primarily form two conflicting requirements; accurate predictions require large context, but precise predictions r

Computer Vision and Pattern RecognitionComputer Science
2
논문|인용수 535·2015
Panoptic Studio: A Massively Multiview System for Social Motion Capture
Hanbyul Joo, Hao Liu, Lei Tan, Lin Gui, Bart Nabbe, Iain Matthews, Takeo Kanade, Shohei Nobuhara, Yaser Sheikh

We present an approach to capture the 3D structure and motion of a group of people engaged in a social interaction. The core challenges in capturing social interactions are: (1) occlusion is functional and frequent, (2) subtle motion needs to be measured over a space large enough to host a social group, and (3) human appearance and configuration variation is immense. The Panoptic Studio is a system organized around the thesis that social interactions should be measured through the perceptual int

Computer Vision and Pattern RecognitionComputer Science
3
논문|인용수 487·2022
Ego4D: Around the World in 3,000 Hours of Egocentric Video
Kristen Grauman, Andrew Westbury, Eugene H. Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, Miguel Martin, Tushar Nagarajan
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)OA

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of dailylife activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countries. The approach to collection is designed to uphold rigorous privacy and ethics standards, with consenting participants and robust de-identification procedures where relevant. Ego4D dramatically expan

Computer Vision and Pattern RecognitionComputer Science
4
논문|인용수 353·2017
Panoptic Studio: A Massively Multiview System for Social Interaction Capture
Hanbyul Joo, Tomas Simon, Xulong Li, Hao Liu, Lei Tan, Lin Gui, Sean Banerjee, Timothy Godisart, Bart Nabbe, Iain Matthews, Takeo Kanade, Shohei Nobuhara
SJR Q1IEEE Transactions on Pattern Analysis and Machine Intelligence

We present an approach to capture the 3D motion of a group of people engaged in a social interaction. The core challenges in capturing social interactions are: (1) occlusion is functional and frequent; (2) subtle motion needs to be measured over a space large enough to host a social group; (3) human appearance and configuration variation is immense; and (4) attaching markers to the body may prime the nature of interactions. The Panoptic Studio is a system organized around the thesis that social

Computer Vision and Pattern RecognitionComputer Science
5
논문|인용수 341·2019
Monocular Total Capture: Posing Face, Body, and Hands in the Wild
Donglai Xiang, Hanbyul Joo, Yaser Sheikh

We present the first method to capture the 3D total motion of a target person from a monocular view input. Given an image or a monocular video, our method reconstructs the motion from body, face, and fingers represented by a 3D deformable mesh model. We use an efficient representation called 3D Part Orientation Fields (POFs), to encode the 3D orientations of all body parts in the common 2D image space. POFs are predicted by a Fully Convolutional Network, along with the joint confidence maps. To

Computer Vision and Pattern RecognitionComputer Science
6
논문|인용수 181·2021
Exemplar Fine-Tuning for 3D Human Model Fitting Towards In-the-Wild 3D Human Pose Estimation
Hanbyul Joo, Natalia Neverova, Andrea Vedaldi
2021 International Conference on 3D Vision (3DV)

Differently from 2D image datasets such as COCO, largescale human datasets with 3D ground-truth annotations are very difficult to obtain in the wild. In this paper, we address this problem by augmenting existing 2D datasets with high-quality 3D pose fits. Remarkably, the resulting annotations are sufficient to train from scratch 3D pose regressor networks that outperform the current state-of-the-art on in the-wild benchmarks such as 3DPW. Additionally, training on our augmented data is straightf

Computer Vision and Pattern RecognitionComputer Science
7
preprint|인용수 140·2017
Hand Keypoint Detection in Single Images Using Multiview Bootstrapping
Tomas Simon, Hanbyul Joo, Iain Matthews, Yaser Sheikh
OA

We present an approach that uses a multi-camera system to train fine-grained detectors for keypoints that are prone to occlusion, such as the joints of a hand. We call this procedure multiview bootstrapping: first, an initial keypoint detector is used to produce noisy labels in multiple views of the hand. The noisy detections are then triangulated in 3D using multiview geometry or marked as outliers. Finally, the reprojected triangulations are used as new labeled training data to improve the det

Computer Vision and Pattern RecognitionComputer Science
8
논문|인용수 108·2022
BANMo: Building Animatable 3D Neural Models from Many Casual Videos
Gengshan Yang, Minh Vo, Natalia Neverova, Deva Ramanan, Andrea Vedaldi, Hanbyul Joo
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

Prior work for articulated 3D shape reconstruction often relies on specialized multi-view and depth sensors or pre-built deformable 3D models. Such methods do not scale to diverse sets of objects in the wild. We present a method that requires neither of them. It aims to create high-fidelity, articulated 3D models from many casual RGB videos in a differentiable rendering framework. Our key in-sight is to merge three schools of thought: (1) classic deformable shape models that make use of articula

Computational MechanicsEngineering
9
book chapter|인용수 106·2020
Perceiving 3D Human-Object Spatial Arrangements from a Single Image in the Wild
Jason Zhang, Sam Pepose, Hanbyul Joo, Deva Ramanan, Jitendra Malik, Angjoo Kanazawa
SJR Q2Lecture notes in computer science
Computational MechanicsEngineering
10
논문|인용수 85·2022
Learning to Listen: Modeling Non-Deterministic Dyadic Facial Motion
Evonne Ng, Hanbyul Joo, Liwen Hu, Hao Li, Trevor Darrell, Angjoo Kanazawa, Shiry Ginosar
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

We present a framework for modeling interactional communication in dyadic conversations: given multimodal inputs of a speaker, we autoregressively output multiple possibilities of corresponding listener motion. We combine the motion and speech audio of the speaker using a motion-audio cross attention transformer. Furthermore, we enable non-deterministic prediction by learning a discrete latent representation of realistic listener motion with a novel motion-encoding VQ-VAE. Our method organically

Signal ProcessingComputer Science
11
논문|인용수 69·2019
Single-Network Whole-Body Pose Estimation
Gines Hidalgo Martinez, Yaadhav Raaj, Haroon Idrees, Donglai Xiang, Hanbyul Joo, Tomas Simon, Yaser Sheikh

We present the first single-network approach for 2D~whole-body pose estimation, which entails simultaneous localization of body, face, hands, and feet keypoints. Due to the bottom-up formulation, our method maintains constant real-time performance regardless of the number of people in the image. The network is trained in a single stage using multi-task learning, through an improved architecture which can handle scale differences between body/foot and face/hand keypoints. Our approach considerabl

Computer Vision and Pattern RecognitionComputer Science
12
preprint|인용수 54·2018
Total Capture: A 3D Deformation Model for Tracking Faces, Hands, and Bodies
Hanbyul Joo, Tomas Simon, Yaser Sheikh
OA

The following topics are dealt with: learning (artificial intelligence); feature extraction; image classification; neural nets; image representation; object detection; image segmentation; convolution; feedforward neural nets; video signal processing.

Computer Vision and Pattern RecognitionComputer Science
13
preprint|인용수 51·2020
Exemplar Fine-Tuning for 3D Human Pose Fitting Towards In-the-Wild 3D Human Pose Estimation
Hanbyul Joo, Natalia Neverova, Andrea Vedaldi
arXiv (Cornell University)OA

We propose a method for building large collections of human poses with full 3D annotations captured `in the wild', for which specialized capture equipment cannot be used. We start with a dataset with 2D keypoint annotations such as COCO and MPII and generates corresponding 3D poses. This is done via Exemplar Fine-Tuning (EFT), a new method to fit a 3D parametric model to 2D keypoints. EFT is accurate and can exploit a data-driven pose prior to resolve the depth reconstruction ambiguity that come

Computer Vision and Pattern RecognitionComputer Science
14
논문|인용수 35·2014
MAP Visibility Estimation for Large-Scale Dynamic 3D Reconstruction
Hanbyul Joo, Hyun Soo Park, Yaser Sheikh

Many traditional challenges in reconstructing 3D motion, such as matching across wide baselines and handling occlusion, reduce in significance as the number of unique viewpoints increases. However, to obtain this benefit, a new challenge arises: estimating precisely which cameras observe which points at each instant in time. We present a maximum a posteriori (MAP) estimate of the time-varying visibility of the target points to reconstruct the 3D motion of an event from a large number of cameras.

Computer Vision and Pattern RecognitionComputer Science
15
preprint|인용수 16·2016
Panoptic Studio: A Massively Multiview System for Social Interaction Capture
Hanbyul Joo, Tomas Simon, Xulong Li, Hao Liu, Lei Tan, Lin Gui, Sean Banerjee, Timothy Godisart, Bart Nabbe, Iain Matthews, Takeo Kanade, Shohei Nobuhara
arXiv (Cornell University)OA

We present an approach to capture the 3D motion of a group of people engaged\nin a social interaction. The core challenges in capturing social interactions\nare: (1) occlusion is functional and frequent; (2) subtle motion needs to be\nmeasured over a space large enough to host a social group; (3) human appearance\nand configuration variation is immense; and (4) attaching markers to the body\nmay prime the nature of interactions. The Panoptic Studio is a system organized\naround the thesis that s

Computer Vision and Pattern RecognitionComputer Science

대표 연구 분야

Computer Vision and Pattern RecognitionControl and Systems EngineeringComputational MechanicsSignal ProcessingBiomedical EngineeringSociology and Political Science

주한별 교수의 연구를 Nubint에서 더 깊이 살펴보세요

이 연구실의 논문을 앱에서 열어 AI와 함께 읽고, 핵심을 요약하고, 내 글에 인용하세요.