Skip to main content
QUICK REVIEW

[논문 리뷰] Clinical-Prior Guided Multi-Modal Learning with Latent Attention Pooling for Gait-Based Scoliosis Screening

Dong Chen, Zizhuang Wei|arXiv (Cornell University)|2026. 02. 06.
Scoliosis diagnosis and treatment인용 수 0
한 줄 요약

ScoliGait를 소개합니다. 비중첩적이며 방사선으로 라벨링된 보행 영상 벤치마크를 AIS 선별에 사용하고, 임상 선행 지식 맵, 비디오, 텍스트를 융합하는 잠재 주의 풀링 기반의 다모달 모델로 해석 가능하고 최첨단 성능을 달성합니다.

ABSTRACT

Adolescent Idiopathic Scoliosis (AIS) is a prevalent spinal deformity whose progression can be mitigated through early detection. Conventional screening methods are often subjective, difficult to scale, and reliant on specialized clinical expertise. Video-based gait analysis offers a promising alternative, but current datasets and methods frequently suffer from data leakage, where performance is inflated by repeated clips from the same individual, or employ oversimplified models that lack clinical interpretability. To address these limitations, we introduce ScoliGait, a new benchmark dataset comprising 1,572 gait video clips for training and 300 fully independent clips for testing. Each clip is annotated with radiographic Cobb angles and descriptive text based on clinical kinematic priors. We propose a multi-modal framework that integrates a clinical-prior-guided kinematic knowledge map for interpretable feature representation, alongside a latent attention pooling mechanism to fuse video, text, and knowledge map modalities. Our method establishes a new state-of-the-art, demonstrating a significant performance gap on a realistic, non-repeating subject benchmark. Our approach establishes a new state of the art, showing a significant performance gain on a realistic, subject-independent benchmark. This work provides a robust, interpretable, and clinically grounded foundation for scalable, non-invasive AIS assessment.

연구 동기 및 목표

  • 데이터 누출과 보행 기반 AIS 선별 데이터셋의 주제 독립성을 해결합니다.
  • 킥매틱 지식 맵을 통해 보행의 임상적으로 근거 있는 해석 가능한 표현을 제공합니다.
  • 비디오, 지식 맵, 텍스트를 잠재 주의 풀링으로 결합하는 강건한 다모달 융합 방법을 개발합니다.

제안 방법

  • 1,572개의 학습 클립과 300개의 독립 테스트 클립으로 구성된 ScoliGait 데이터셋을 제안하고, 각 클립은 방사선 Cobb 각도와 임상 텍스트 프롬프트로 주석을 달았습니다.
  • 모션 공간, 자기 골격 공간, 신호 상관관계에 걸친 238개의 특징으로 구성된 운동학 지식 맵을 구축합니다.
  • 지식 맵, 비디오, 텍스트의 세 가지 모듈별 인코더와 잠재 주의 풀링 메커니즘을 사용하여 모달리티를 융합합니다.
  • 융합 성능을 향상시키기 위해 모달 간 위치 임베딩을 정렬합니다.
  • 임상 지식 맵으로 주의 점수를 매핑하여 해석 가능성을 제공합니다.
  • 텍스트 인코딩에 Sentence-Transformers를, 비디오 및 지식 맵 모달리티에 Vision Transformer 백본을 적용합니다.
Figure 1: ScoliGait system for multi-modal gait analysis from mobile video. Left: temporal alignment of the knowledge map and video. Right: generation of video, knowledge map, and text modalities via pose estimation, showing kinematic alignment and knowledge-guided synthesis.
Figure 1: ScoliGait system for multi-modal gait analysis from mobile video. Left: temporal alignment of the knowledge map and video. Right: generation of video, knowledge map, and text modalities via pose estimation, showing kinematic alignment and knowledge-guided synthesis.

실험 결과

연구 질문

  • RQ1임상적으로 근거가 있는 다모달 프레임워크가 주제 독립 보행 데이터세트에서 AIS 선별 정확도를 향상시킬 수 있는가?
  • RQ2구조화된 운동학 지식 맵을 비디오 및 텍스트와 결합하는 것이 해석 가능성과 척추 기형 선별의 진단 성능을 향상시키는가?
  • RQ3잠재 주의 풀링과 교차 모달 정렬이 융합 품질 및 임상 관련성에 미치는 영향은 무엇인가?

주요 결과

  • 지식 맵만으로도 이진 AIS 선별에서 비디오만으로의 정확도 대비 1.7%의 향상과 F1-score 3.2%의 향상을 보인다.
  • 지식 맵, 비디오, 텍스트의 잠재 주의 풀링을 이용한 다모달 융합이 최상의 성능을 달성하며: 정확도 70.0%와 F1-score 61.9%이다.
  • ScoliGait은 고유한 개인들로부터 1,572개의 학습 클립과 300개의 독립 테스트 클립을 제공하는 비중첩적이면서 방사선으로 라벨링된 벤치마크를 제공한다.
  • 주의를 임상적으로 의미 있는 지식 맵으로 매핑하여 보행 특징을 시간에 따라 명시적으로 해석 가능하게 함으로써 해석 가능성이 향상된다.
  • 아블레이션 결과, 잠재 주의 풀링이 단순 연결(concatenation)보다 우수하고 교차 모달 임베딩 정렬이 결과를 향상시킨다.
Figure 2: Proposed three-modal fusion architecture for AIS screening. Inputs from Knowledge Map, Vision, and Text modalities are integrated via a Latent Attention Pooling mechanism (bottom). Remapped attention scores from the Knowledge Map (top) are filtered for salient features to enable clinical i
Figure 2: Proposed three-modal fusion architecture for AIS screening. Inputs from Knowledge Map, Vision, and Text modalities are integrated via a Latent Attention Pooling mechanism (bottom). Remapped attention scores from the Knowledge Map (top) are filtered for salient features to enable clinical i

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.