Skip to main content
QUICK REVIEW

[논문 리뷰] Motion-to-Response Content Generation via Multi-Agent AI System with Real-Time Safety Verification

HyeYoung Lee|arXiv (Cornell University)|2026. 01. 20.
Emotion and Mood Recognition인용 수 0
한 줄 요약

이 논문은 오디오 기반 감정을 안전하고 실시간, 연령에 적합한 응답 콘텐츠로 변환하는 네 에이전트 시스템과 안전 검증 루프 및 디바이스 내 배치를 제시합니다.

ABSTRACT

This paper proposes a multi-agent artificial intelligence system that generates response-oriented media content in real time based on audio-derived emotional signals. Unlike conventional speech emotion recognition studies that focus primarily on classification accuracy, our approach emphasizes the transformation of inferred emotional states into safe, age-appropriate, and controllable response content through a structured pipeline of specialized AI agents. The proposed system comprises four cooperative agents: (1) an Emotion Recognition Agent with CNN-based acoustic feature extraction, (2) a Response Policy Decision Agent for mapping emotions to response modes, (3) a Content Parameter Generation Agent for producing media control parameters, and (4) a Safety Verification Agent enforcing age-appropriateness and stimulation constraints. We introduce an explicit safety verification loop that filters generated content before output, ensuring compliance with predefined rules. Experimental results on public datasets demonstrate that the system achieves 73.2% emotion recognition accuracy, 89.4% response mode consistency, and 100% safety compliance while maintaining sub-100ms inference latency suitable for on-device deployment. The modular architecture enables interpretability and extensibility, making it applicable to child-adjacent media, therapeutic applications, and emotionally responsive smart devices.

연구 동기 및 목표

  • 감정 인식과 콘텐츠 생성을 명시적 정책 및 안전 계층과 연계한다.
  • 해석 가능하고 모듈식이며 디바이스 내 감정-응답 콘텐츠 생성을 가능하게 한다.
  • 규칙 기반 안전 검증을 통해 연령 적합성 및 통제된 자극을 보장한다.
  • 실시간 성능과 프라이버시 보존, 엣지 중심 배치를 입증한다.

제안 방법

  • 네 가지 협력 에이전트가 입력을 처리한다: Emotion Recognition, Response Policy Decision, Content Parameter Generation, 및 Safety Verification.
  • CNN 기반 음향 특징 추출 및 softmax 기반 감정 분류를 e* in C 감정 범주에서.
  • 감정과 각성에서 이산 응답 모드로의 정책 매핑을 의사 결정 트리를 통해 수행한다.
  • 선택된 모드에서 다중 모달 미디어 제어(오디오, 비주얼, 텍스트)를 예측하는 콘텐츠 매개 변수 생성.
  • 규칙 기반 제약을 사용한 명시적 안전 검증으로, 규칙이 위반될 경우 재생성 루프를 포함한다.
Figure 1: Overall architecture of the proposed multi-agent system for emotion-to-response content generation. The system comprises four specialized agents operating sequentially: Emotion Recognition Agent, Response Policy Decision Agent, Content Parameter Generation Agent, and Safety Verification Ag
Figure 1: Overall architecture of the proposed multi-agent system for emotion-to-response content generation. The system comprises four specialized agents operating sequentially: Emotion Recognition Agent, Response Policy Decision Agent, Content Parameter Generation Agent, and Safety Verification Ag

실험 결과

연구 질문

  • RQ1경량의 on-device 설정에서 오디오로부터 감정을 얼마나 정확하게 인식할 수 있는가?
  • RQ2인식된 감정이 안전하고 연령에 적합한 응답 모드로 신뢰성 있게 매핑될 수 있는가?
  • RQ3생성된 콘텐츠 매개 변수가 실시간 검증으로 안전 제약을 충족하는가?
  • RQ4서로 다른 하드웨어에서 시스템의 엔드투엔드 대기시간은 어느 정도이며, on-device 배치에 적합한가?
  • RQ5안전 검증 루프가 출력 품질과 신뢰성에 미치는 영향은 무엇인가?

주요 결과

  • 감정 인식 정확도는 데이터 세트에 따라 달라지며, 4클래스에서 IEMOCAP 73.2%, 4클래스에서 RAVDESS 78.5%, 합성 데이터에서 89.3%입니다.
  • 응답 모드 정확도는 89.4%이며, 매크로 지표에서 정밀도와 재현율이 대략 86–87% 수준입니다.
  • 규칙 전반에 걸쳐 안전 검증이 100% 합격률을 달성했고, 필요한 재생성은 단 1.6%에 불과했습니다.
  • 평가된 하드웨어에서 엔드투엔드 추론 대기시간이 100 ms 미만으로 유지되었으며, 엣지 디바이스를 포함합니다.
  • 제거 분석은 다에이전트 구조가 엔드투엔드 구성에 비해 일관성과 안전성을 향상시킴을 보여줍니다.
Figure 2: Processing flowchart of the emotion-to-response content generation method. The system processes audio input through sequential stages (S10-S60) with a safety verification loop that triggers content regeneration upon failure.
Figure 2: Processing flowchart of the emotion-to-response content generation method. The system processes audio input through sequential stages (S10-S60) with a safety verification loop that triggers content regeneration upon failure.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.