[論文レビュー] LaERC-S: Improving LLM-based Emotion Recognition in Conversation with Speaker Characteristics
LaERC-Sは、歴史的な発話を用いて話者の常識を生成し、それを用いて ERC の性能を向上させることで、会話における感情認識を強化します。
Emotion recognition in conversation (ERC), the task of discerning human emotions for each utterance within a conversation, has garnered significant attention in human-computer interaction systems. Previous ERC studies focus on speaker-specific information that predominantly stems from relationships among utterances, which lacks sufficient information around conversations. Recent research in ERC has sought to exploit pre-trained large language models (LLMs) with speaker modelling to comprehend emotional states. Although these methods have achieved encouraging results, the extracted speaker-specific information struggles to indicate emotional dynamics. In this paper, motivated by the fact that speaker characteristics play a crucial role and LLMs have rich world knowledge, we present LaERC-S, a novel framework that stimulates LLMs to explore speaker characteristics involving the mental state and behavior of interlocutors, for accurate emotion predictions. To endow LLMs with this knowledge information, we adopt the two-stage learning to make the models reason speaker characteristics and track the emotion of the speaker in complex conversation scenarios. Extensive experiments on three benchmark datasets demonstrate the superiority of LaERC-S, reaching the new state-of-the-art.
研究の動機と目的
- 話者に関連する常識とリスナーの反応を取り入れることで、ERCの改善を促進する。
- 歴史的発話を用いて話者中心の常識を大規模言語モデルで生成する。
- 話者の常識識別で事前学習を行い、ERC性能を向上させる。
- 強化された話者コンテクスト機能を備えたLLMベースのERCシステムをファインチューニングする。
- 標準的なERCデータセットで最先端または競争力のある結果を示す。
提案手法
- 歴史的発話を用いて、意図やリスナーの反応に関する話者の常識を Llama2-Chat プロンプトを通じて生成する。
- 9つの ATOMIC 関係(xIntent, xReact, oReact など)を標的とした関係ベースのテンプレートを構築して常識データを生成する。
- 事前学習のために話者識別の補助タスクを対話相手の常識識別に置き換える。
- パラメータを制御するために LoRA を用いた Llama2 系モデルを用いた2段階プロセスで ERC モデルをファインチューニングする。
- 常識識別と感情予測の両方に対して、履歴・タスク定義・期待出力を含むプロンプトを活用する。
- IEMOCAP、EmoryNLP、MELDで重み付き F1 を主要指標として評価する。
実験結果
リサーチクエスチョン
- RQ1歴史的対話からの対話者常識を取り入れることで、話者識別のベースラインよりERCの精度を改善できるか?
- RQ2リスナーの反応と話者の意図を生成するようにLLMプロンプトを調整すると、感情認識が向上するか?
- RQ3InstructERC および他の常識ベースのモデルと比較して、LaERC-S は標準的なERCデータセットでどの程度の性能を示すか?
主な発見
- LaERC-Sは IEMOCAP、EmoryNLP、MELD の全てで最先端または競争力のある結果を達成します。
- 平均的に、LaERC-S は InstructERC および複数のベースラインを上回り、特に IEMOCAP と EmoryNLP データセットで顕著な向上を示します。
- 歴史的ディスコースを用いて現在の発話常識を生成することで、トークンレベルやディスコース非依存の方法よりもより正確な暗黙の感情手掛かりを得られます。
- ERC 本タスクの前に対話相手の常識識別を用いた二段階トレーニングは感情予測を改善します。
- 3データセットにまたがる平均結果は、選択したベースラインに対して有利な性能向上を示します。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。