Skip to main content
QUICK REVIEW

[论文解读] Memory Bear AI Memory Science Engine for Multimodal Affective Intelligence: A Technical Report

Deliang Wen, Ke Sun|arXiv (Cornell University)|Mar 18, 2026
Emotion and Mood Recognition被引用 0
一句话总结

Memory Bear AI Memory Science Engine 将情感信息建模为结构化记忆系统(EMUs),以实现长时域、鲁棒的多模态情感判断与检索,在多个数据集和嘈杂条件下优于基线。

ABSTRACT

Affective judgment in real interaction is rarely a purely local prediction problem. Emotional meaning often depends on prior trajectory, accumulated context, and multimodal evidence that may be weak, noisy, or incomplete at the current moment. Although multimodal emotion recognition (MER) has improved the integration of text, speech, and visual signals, many existing systems remain optimized for short-range inference and provide limited support for persistent affective memory, long-horizon dependency modeling, and robust interpretation under imperfect input. This technical report presents the Memory Bear AI Memory Science Engine, a memory-centered framework for multimodal affective intelligence. Instead of treating emotion as a transient output label, the framework models affective information as a structured and evolving variable within a memory system. It organizes processing through structured memory formation, working-memory aggregation, long-term consolidation, memory-driven retrieval, dynamic fusion calibration, and continuous memory updating. At its core, multimodal signals are transformed into structured Emotion Memory Units (EMUs), enabling affective information to be preserved, reactivated, and revised across interaction horizons. Experimental results show consistent gains over comparison systems across benchmark and business-grounded settings, with stronger accuracy and robustness, especially under noisy or missing-modality conditions. The framework offers a practical step from local emotion recognition toward more continuous, robust, and deployment-relevant affective intelligence.

研究动机与目标

  • 将情感判断重新表述为以记忆为中心的问题,而非纯粹的局部预测任务。
  • 提出一种结构化记忆体系,将多模态证据编码为可复用的情感记忆单元(EMUs)。
  • 实现基于记忆的短期与长期聚合、检索与动态融合,以在缺失或嘈杂模态下提高鲁棒性。
  • 在基准数据集和面向商业的数据集上展示更强的性能与鲁棒性,并给出面向部署的分析。

提出的方法

  • 阶段1:多模态预处理与表示学习,生成模态特定的情感编码(文本通过基于大模型的语义编码;音频通过 Higgs-Audio;视觉通过 VLM 驱动的表示)。
  • 阶段2:结构化情感记忆建模,形成捕捉情感 e_t、源可靠性 m_t、情境锚 c_t、显著性 α_t、时间 τ_t 的 EMU。
  • 阶段2 还包括情感工作记忆以进行短期聚合,以及情感长期记忆以实现巩固,还有基于记忆的检索。
  • 阶段3:动态融合策略,将多模态贡献与历史记忆进行标定。
  • 阶段4:分类、决策制定与记忆更新,记忆生命周期包括遗忘与更新。

实验结果

研究问题

  • RQ1记忆中心设计在长互动时域中如何影响情感判断的稳定性与准确性?
  • RQ2与传统融合方法相比,EMUs与记忆驱动的检索是否能在模态缺失或降级时提升鲁棒性?
  • RQ3在标准 MER 基准(IEMOCAP、CMU-MOSEI)与面向商业的数据集上,准确率与稳定性有何提升?
  • RQ4记忆引导的标定如何在嘈杂输入下影响实时情感解读?

主要发现

  • 在 IEMOCAP 上,Memory Bear AI 的准确率为 78.8%。
  • 在 CMU-MOSEI 上,Memory Bear AI 的准确率为 66.7%。
  • 在 Memory Bear AI Business Dataset 上,准确率为 68.4%,加权 F1 为 48.6,宏 F1 为 45.9%。
  • 该模型在商业数据集上相较传统融合基线展现出更强的精度提升(8.2 点)。
  • 在降级多模态条件下,该框架保持了完整条件性能的 92.3%,显示鲁棒性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。