[论文解读] Clinical-Prior Guided Multi-Modal Learning with Latent Attention Pooling for Gait-Based Scoliosis Screening
介绍 ScoliGait:一个用于 AIS 筛查的非重叠、放射标注的步态视频基准,以及一个具有潜在注意力汇聚的临床先验引导多模态模型,将知识图、视频和文本融合作为可解释的、最先进性能的系统。
Adolescent Idiopathic Scoliosis (AIS) is a prevalent spinal deformity whose progression can be mitigated through early detection. Conventional screening methods are often subjective, difficult to scale, and reliant on specialized clinical expertise. Video-based gait analysis offers a promising alternative, but current datasets and methods frequently suffer from data leakage, where performance is inflated by repeated clips from the same individual, or employ oversimplified models that lack clinical interpretability. To address these limitations, we introduce ScoliGait, a new benchmark dataset comprising 1,572 gait video clips for training and 300 fully independent clips for testing. Each clip is annotated with radiographic Cobb angles and descriptive text based on clinical kinematic priors. We propose a multi-modal framework that integrates a clinical-prior-guided kinematic knowledge map for interpretable feature representation, alongside a latent attention pooling mechanism to fuse video, text, and knowledge map modalities. Our method establishes a new state-of-the-art, demonstrating a significant performance gap on a realistic, non-repeating subject benchmark. Our approach establishes a new state of the art, showing a significant performance gain on a realistic, subject-independent benchmark. This work provides a robust, interpretable, and clinically grounded foundation for scalable, non-invasive AIS assessment.
研究动机与目标
- 解决步态基于 AIS 筛查数据集中数据泄露和受试者独立性问题。
- 通过运动学知识图提供一个临床、可解释的步态表示。
- 开发一个鲁棒的多模态融合方法,将视频、知识图和文本与潜在注意力汇聚相结合。
提出的方法
- 提出具有 1,572 条训练片段和 300 条独立测试片段的 ScoliGait 数据集,每条片段均标注放射 Cobb 角和临床文本提示。
- 构建包含运动空间、自骨架空间和信号相关性的 238 个特征的运动学知识图。
- 使用三个模态专用的编码器(知识图、视频、文本)及潜在注意力汇聚机制来融合模态。
- 跨模态对齐位置嵌入以提升融合性能。
- 通过将注意力分数映射回临床知识图来提供可解释的解释。
- 为文本编码采用 Sentence-Transformers,为视频和知识图模态采用 Vision Transformer 主干网络。

实验结果
研究问题
- RQ1一个以临床为基础的多模态框架是否能够在一个主体独立的步态数据集上改善 AIS 筛查的准确性?
- RQ2将结构化的运动学知识图与视频和文本结合是否能提升对脊柱侧弯筛查的可解释性和诊断性能?
- RQ3潜在注意力汇聚和跨模态对齐对融合质量及临床相关性的影响如何?
主要发现
- 知识图单独在二分类 AIS 筛查中的准确率比单独的视频高出 1.7%、F1-score 高出 3.2%。
- 结合知识图、视频与文本并使用潜在注意力汇聚的多模态融合达到最佳性能:准确率 70.0%,F1-score 61.9%。
- ScoliGait 提供一个非重叠、放射标注的基准,包含 1,572 条训练片段和来自不同个体的 300 条独立测试片段。
- 通过将注意力映射到临床有意义的知识图上,提升了对步态特征随时间的可解释性。
- 消融研究显示潜在注意力汇聚优于简单拼接,且跨模态嵌入对齐可改善结果。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。