Skip to main content
QUICK REVIEW

[논문 리뷰] 3DPalsyNet: A Facial Palsy Grading and Motion Recognition Framework using Fully 3D Convolutional Neural Networks

Gary Storey, Richard Jiang|arXiv (Cornell University)|2019. 05. 31.
Facial Nerve Paralysis Treatment and Research참고 문헌 27인용 수 4
한 줄 요약

3DPalsyNet는 3D 컨volution 신경망 아키텍처에 ResNet 기반을 채택한 프레임워크로, 비디오 시퀀스에서 얼굴마비 정도 평가 및 입동작 인식을 엔드 투 엔드로 수행할 수 있다. Kinetics에서의 전이학습을 활용하고, 중심손실(Center Loss)과 소프트맥스 손실(Softmax Loss)을 조합함으로써, 8프레임 지속 시간에서 운동 인식 정확도 86%와 평가 정확도 82%를 달성하였으며, 최적의 성능를 기록하였다.

ABSTRACT

The capability to perform facial analysis from video sequences has significant potential to positively impact in many areas of life. One such area relates to the medical domain to specifically aid in the diagnosis and rehabilitation of patients with facial palsy. With this application in mind, this paper presents an end-to-end framework, named 3DPalsyNet, for the tasks of mouth motion recognition and facial palsy grading. 3DPalsyNet utilizes a 3D CNN architecture with a ResNet backbone for the prediction of these dynamic tasks. Leveraging transfer learning from a 3D CNNs pre-trained on the Kinetics data set for general action recognition, the model is modified to apply joint supervised learning using center and softmax loss concepts. 3DPalsyNet is evaluated on a test set consisting of individuals with varying ranges of facial palsy and mouth motions and the results have shown an attractive level of classification accuracy in these task of 82% and 86% respectively. The frame duration and the loss function affect was studied in terms of the predictive qualities of the proposed 3DPalsyNet, where it was found shorter frame duration's of 8 performed best for this specific task. Centre loss and softmax have shown improvements in spatio-temporal feature learning than softmax loss alone, this is in agreement with earlier work involving the spatial domain.

연구 동기 및 목표

  • 비디오 데이터를 활용한 자동 얼굴마비 평가 및 입운동 인식을 위한 엔드 투 엔드 프레임워크 개발
  • 중심손실과 소프트맥스 손실을 병행 적용한 공동 학습을 통해 얼굴운동 분석에서의 시공간 특징 학습 향상
  • 프레임 지속 시간과 손실 함수 설계가 얼굴마비 평가 분류 성능에 미치는 영향 평가
  • 강력한 딥러닝 기반 영상 분석을 통해 진단 및 재활 지원 분야의 임상 적용 가능화
  • 의료적 맥락에서 동적 얼굴표정 분석에 3D CNN과 ResNet 아키텍처의 효과성 입증

제안 방법

  • 프레임워크는 비디오 클립의 얼굴 운동에서 시공간 특징을 추출하기 위해 3D CNN과 ResNet 기반 백본을 활용한다.
  • 일반적인 동작 인식을 위한 3D CNN을 Kinetics 데이터셋에서 미리 학습시킨 결과를 활용해 전이학습을 적용한다.
  • 특징의 분리 능력을 향상시키기 위해 중심손실과 소프트맥스 손실을 모두 사용하는 공동 학습을 구현한다.
  • 다양한 수준의 얼굴마비와 운동 유형을 가진 개인으로 구성된 다양한 테스트 세트에서 모델을 학습하고 평가한다.
  • 예측 성능에 미치는 영향을 평가하기 위해 지속 시간을 체계적으로 변화시켜(예: 8, 16, 32 프레임) 분석한다.
  • 얼굴운동 시퀀스의 순차적 동역학을 유지함으로써 시간 모델링에 최적화된 아키텍처를 설계한다.

실험 결과

연구 질문

  • RQ1ResNet 기반 3D CNN은 얼굴마비와 관련된 동적 얼굴운동을 인식하는 데 얼마나 효과적인가?
  • RQ2얼굴마비 평가 및 운동 인식에서 정확한 분류를 위한 최적의 프레임 지속 시간은 무엇인가?
  • RQ3소프트맥스 손실 단독 사용 대비 중심손실과 소프트맥스 손실의 조합이 시공간 특징 학습에 어떤 영향을 미치는가?
  • RQ4Kinetics에서의 전이학습은 얼굴마비 분류 작업 성능 향상에 기여하는가?
  • RQ5다양한 수준의 얼굴마비 심각도를 가진 환자들 간에 모델의 일반화 능력은 어떠한가?

주요 결과

  • 3DPalsyNet 프레임워크는 테스트 세트에서 입운동 인식 분류 정확도 86%를 달성하였다.
  • 제안된 프레임워크를 사용하여 얼굴마비 평가 정확도는 82%를 기록하였다.
  • 8프레임의 짧은 지속 시간에서 가장 높은 성능를 기록하였으며, 16 또는 32프레임보다 뛰어난 성능를 보였다.
  • 중심손실과 소프트맥스 손실의 조합은 소프트맥스 손실 단독 사용 대비 시공간 특징 학습을 크게 향상시켰다.
  • Kinetics 데이터셋에서의 전이학습을 통해 모델의 일반화 능력과 얼굴마비 작업 성능이 향상되었다.
  • 결과는 3D CNN을 활용한 영상 기반 자동으로 임상적으로 유의미한 얼굴마비 평가의 가능성을 입증하였다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.