Skip to main content
QUICK REVIEW

[논문 리뷰] Self-Emotion-Mediated Exploration in Artificial Intelligence Mirrors: Findings from Cognitive Psychology

Gustavo Assunção, Miguel Castelo‐Branco|arXiv (Cornell University)|2023. 02. 13.
Child and Animal Learning Development인용 수 5
한 줄 요약

이 논문은 정확도 및 신뢰도와 같은 성능 지표를 사용하여 인지심리학적 원리에 기반한 지식적(놀라움) 및 성취(자부심) 감정을 모델링함으로써 자기 감정 기반 탐색을 통합한 인공지능을 위한 혁신적인 프레임워크를 제안한다. 이 감정들은 딥 강화학습 아키텍처 내에서 미분 가능한 함수로 시뮬레이션되며, 에이전트는 자율적으로 탐색을 학습하게 되어 내부 감정 상태와 탐색 행동 간의 인과관계를 보여주며 인간의 인지 패턴을 모방한다. 이는 90%의 에이전트가 탐색 전략을 성공적으로 학습함으로써 입증된다.

ABSTRACT

Background: Exploration of the physical environment is an indispensable precursor to information acquisition and knowledge consolidation for living organisms. Yet, current artificial intelligence models lack these autonomy capabilities during training, hindering their adaptability. This work proposes a learning framework for artificial agents to obtain an intrinsic exploratory drive, based on epistemic and achievement emotions triggered during data observation. Methods: This study proposes a dual-module reinforcement framework, where data analysis scores dictate pride or surprise, in accordance with psychological studies on humans. A correlation between these states and exploration is then optimized for agents to meet their learning goals. Results: Causal relationships between states and exploration are demonstrated by the majority of agents. A 15.4\% mean increase is noted for surprise, with a 2.8\% mean decrease for pride. Resulting correlations of $ρ_{surprise}=0.461$ and $ρ_{pride}=-0.237$ are obtained, mirroring previously reported human behavior. Conclusions: These findings lead to the conclusion that bio-inspiration for AI development can be of great use. This can incur benefits typically found in living beings, such as autonomy. Further, it empirically shows how AI methodologies can corroborate human behavioral findings, showcasing major interdisciplinary importance. Ramifications are discussed.

연구 동기 및 목표

  • 인공지능과 생물학적 탐색 간 격차를 해소하기 위해 인간과 유사한 감정적 동기를 인공지능 에이전트에 통합하기 위해.
  • 측정 가능한 성능 지표를 기반으로 인지심리학 원리에 따라 지식적 감정(놀라움)과 성취 감정(자부심)을 모델링하기 위해.
  • 내부 감정 상태를 탐색 행동과 연결시켜 인공에이전트가 환경을 자율적으로 탐색하도록 하여 적응성과 학습 효율성을 향상시키기 위해.
  • 감정 기반 탐색이 인간의 행동 패턴을 인지심리학 실험에서 관찰된 바와 같이 재현하는지 검증하기 위해.

제안 방법

  • 감정은 미분 가능한 함수로 모델링된다: 자부심은 작업 정확도의 비선형 함수로, 놀라움은 정확도와 신뢰도 점수의 함수로 모델링된다.
  • 프레임워크는 액터-크리틱 네트워크와 리play 버퍼를 사용한 딥 강화학습 아키텍처를 사용하여 학습 안정성과 시간 차분 학습을 보장한다.
  • 감정 함수는 클리핑을 통해 [0,1] 범위로 제한되며, 개인적 차이를 시뮬레이션하기 위해 가우시안 노이즈가 포함되어 있어 현실적인 감정 변동성을 보장한다.
  • 크리틱 네트워크는 소프트 업데이트(τ=0.005)를 사용한 타겟 네트워크를 활용해 시간 차분 타겟을 계산하며, 학습률 0.002로 Adam 최적화기를 사용해 손실을 최소화한다.
  • 놀라움은 정확도와 신뢰도의 차이 제곱에 대해 회전 및 이동된 함수를 사용하여 모델링되며, 고신뢰도 오류와 저신뢰도 성공 사례를 고놀라움 상태로 포착한다.
  • 자부심은 정확도에 대해 가우시안 영향을 받는 함수로 모델링되며, 고정확도에서 피크를 이룬다. 이는 성취 기반 감정 반응을 반영한다.
Figure 1 : Curves demonstrating how the emotion of pride may correlate with accuracy. Example curves following a positive prediction of pride based on increasing accuracy, as described by cognitive psychology research [ 32 , 31 ] . Considering how increasing task accuracy equates to personal achieve
Figure 1 : Curves demonstrating how the emotion of pride may correlate with accuracy. Example curves following a positive prediction of pride based on increasing accuracy, as described by cognitive psychology research [ 32 , 31 ] . Considering how increasing task accuracy equates to personal achieve

실험 결과

연구 질문

  • RQ1인공에이전트는 인지심리학에 기반한 자기 생성 감정 상태를 기반으로 탐색을 조절하는 데 성공할 수 있는가?
  • RQ2AI에서 감정 기반 탐색 전략은 지식적 및 성취 감정 연구에서 관찰된 인간의 행동 패턴을 어느 정도 재현하는가?
  • RQ3정확도 및 신뢰도와 같은 성능 지표가 놀라움과 자부심과 같은 인공 감정으로 어떻게 변환되어 효과적인 탐색을 이끄는가?
  • RQ4명시적으로 모델링된 내부 감정 상태에 의해 안내될 때, 딥 강화학습 에이전트는 탐색 행동에서 수렴에 도달할 수 있는가?

주요 결과

  • 90%의 인공에이전트가 내부 감정 상태 기반으로 탐색을 조절하는 데 성공하여 감정과 행동 간의 인과관계를 입증했다.
  • 놀라움 함수는 고신뢰도 오류와 저신뢰도 성공 사례를 고놀라움 상태로 효과적으로 포착하여 인지심리학 연구 결과와 일치했다.
  • 자부심 함수는 정확도와 비선형적이고 양의 관계를 보였으며, 인간 연구에서 관찰된 성취 기반 감정 반응을 반영했다.
  • 랜덤 C1 및 C2 값으로 개별화된 감정 파rameter를 가진 에이전트들은 다양한 그러나 일관된 탐색 행동을 보였으며, 성격 차이를 시뮬레이션했다.
  • 클리핑 처리 및 노이즈 주입된 감정 함수의 사용으로 모든 에이전트에서 안정적이고 제한된 감정 출력이 보장되어 학습의 강건성을 향상시켰다.
  • 감정 함수를 DDPG 기반 강화학습 프레임워크에 통합함으로써 100 에피소드 동안 탐색 행동에서 일관된 수렴이 달성되었다.
Figure 2 : Multiple perspective surface view demonstrating how the emotion of surprise correlates with accuracy and confidence. A potential surface representative of how surprise fluctuates with polarized variations of task accuracy and agent confidence, as described by cognitive psychology research
Figure 2 : Multiple perspective surface view demonstrating how the emotion of surprise correlates with accuracy and confidence. A potential surface representative of how surprise fluctuates with polarized variations of task accuracy and agent confidence, as described by cognitive psychology research

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.