Skip to main content
QUICK REVIEW

[Paper Review] Memory Bear AI Memory Science Engine for Multimodal Affective Intelligence: A Technical Report

Deliang Wen, Ke Sun|arXiv (Cornell University)|Mar 18, 2026
Emotion and Mood Recognition0 citations
TL;DR

The Memory Bear AI Memory Science Engine models affective information as a structured memory system (EMUs) to enable long-horizon, robust multimodal affective judgment and retrieval, outperforming baselines on several datasets and under noisy conditions.

ABSTRACT

Affective judgment in real interaction is rarely a purely local prediction problem. Emotional meaning often depends on prior trajectory, accumulated context, and multimodal evidence that may be weak, noisy, or incomplete at the current moment. Although multimodal emotion recognition (MER) has improved the integration of text, speech, and visual signals, many existing systems remain optimized for short-range inference and provide limited support for persistent affective memory, long-horizon dependency modeling, and robust interpretation under imperfect input. This technical report presents the Memory Bear AI Memory Science Engine, a memory-centered framework for multimodal affective intelligence. Instead of treating emotion as a transient output label, the framework models affective information as a structured and evolving variable within a memory system. It organizes processing through structured memory formation, working-memory aggregation, long-term consolidation, memory-driven retrieval, dynamic fusion calibration, and continuous memory updating. At its core, multimodal signals are transformed into structured Emotion Memory Units (EMUs), enabling affective information to be preserved, reactivated, and revised across interaction horizons. Experimental results show consistent gains over comparison systems across benchmark and business-grounded settings, with stronger accuracy and robustness, especially under noisy or missing-modality conditions. The framework offers a practical step from local emotion recognition toward more continuous, robust, and deployment-relevant affective intelligence.

Motivation & Objective

  • Reframe affective judgment as a memory-centered problem rather than a purely local prediction task.
  • Propose a structured memory architecture that encodes multimodal evidence into reusable Emotion Memory Units (EMUs).
  • Enable memory-based short-term and long-term aggregation, retrieval, and dynamic fusion to improve robustness under missing or noisy modalities.
  • Demonstrate stronger performance and robustness on benchmark datasets and a business-focused dataset, with deployment-oriented analysis.

Proposed method

  • Stage 1: Multimodal preprocessing and representation learning to produce modality-specific affective encodings (text via LLM-based semantic encoding; audio via Higgs-Audio; vision via VLM-driven representations).
  • Stage 2: Structured affective memory modeling that forms EMUs capturing emotion e_t, source reliability m_t, contextual anchor c_t, salience α_t, and temporal τ_t.
  • Stage 2 also includes emotion working memory for short-term aggregation and emotion long-term memory for consolidation, plus memory-driven retrieval.
  • Stage 3: Dynamic fusion strategies that calibrate multimodal contributions against historical memory.
  • Stage 4: Classification, decision-making, and memory updating with a memory lifecycle that includes forgetting and updating.

Experimental results

Research questions

  • RQ1How does a memory-centered design affect stability and accuracy of affective judgments over long interaction horizons?
  • RQ2Can EMUs and memory-driven retrieval improve robustness under missing or degraded modalities compared to traditional fusion approaches?
  • RQ3What are the gains in accuracy and stability on standard MER benchmarks (IEMOCAP, CMU-MOSEI) and business-oriented datasets?
  • RQ4How does memory-guided calibration influence real-time affective interpretation under noisy inputs?

Key findings

  • On IEMOCAP, Memory Bear AI achieves 78.8% accuracy.
  • On CMU-MOSEI, Memory Bear AI achieves 66.7% accuracy.
  • On Memory Bear AI Business Dataset, accuracy is 68.4%, weighted F1 is 48.6, and macro F1 is 45.9.
  • The model reproduces stronger accuracy gains (8.2 points) over a traditional fusion baseline on the Business Dataset.
  • Under degraded multimodal conditions, the framework preserves 92.3% of complete-condition performance, showing robustness.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.