Skip to main content
QUICK REVIEW

[論文レビュー] InvestLM: A Large Language Model for Investment using Financial Domain Instruction Tuning

Yi Yang, Yixuan Tang|arXiv (Cornell University)|Sep 15, 2023
Stock Market Forecasting Methods被引用数 24
ひとこと要約

InvestLMは、厳選された、ドメイン特化の指示データセットを用いてLLaMA-65Bを指示学習させた金融分野のLLMであり、専門家評価で競争力を持ち、金融NLPベンチマークにおけるタスク横断的な一般化能力が高い。

ABSTRACT

We present a new financial domain large language model, InvestLM, tuned on LLaMA-65B (Touvron et al., 2023), using a carefully curated instruction dataset related to financial investment. Inspired by less-is-more-for-alignment (Zhou et al., 2023), we manually curate a small yet diverse instruction dataset, covering a wide range of financial related topics, from Chartered Financial Analyst (CFA) exam questions to SEC filings to Stackexchange quantitative finance discussions. InvestLM shows strong capabilities in understanding financial text and provides helpful responses to investment related questions. Financial experts, including hedge fund managers and research analysts, rate InvestLM's response as comparable to those of state-of-the-art commercial models (GPT-3.5, GPT-4 and Claude-2). Zero-shot evaluation on a set of financial NLP benchmarks demonstrates strong generalizability. From a research perspective, this work suggests that a high-quality domain specific LLM can be tuned using a small set of carefully curated instructions on a well-trained foundation model, which is consistent with the Superficial Alignment Hypothesis (Zhou et al., 2023). From a practical perspective, this work develops a state-of-the-art financial domain LLM with superior capability in understanding financial texts and providing helpful investment advice, potentially enhancing the work efficiency of financial professionals. We release the model parameters to the research community.

研究の動機と目的

  • 投資タスクのためのアクセス可能で高品質な金融分野LLMの必要性を動機付ける。
  • 小規模で入念に厳選されたドメイン指示セットが、基盤モデルを効果的にファインチューンできることを示す。
  • ドメイン指示チューニングが金融NLPベンチマークへの強い一般化をもたらすことを示す。
  • InvestLMを最先端の商用モデルと比較した専門家評価を提供する。
  • ドメイン特異的データと汎用指示データがモデルの性能に与える影響について洞察を提供する。

提案手法

  • LoRA(rank 16)を用いてLLaMA-65Bをファインチューンする。
  • 長い金融テキスト向けにContext長を8,192トークンへ拡張するLinear Rope Scalingを使用する。
  • InvestLM-65Bで学習率3e-4、バッチサイズ16で15エポック訓練する。
  • CFA、StackExchange QFin、学術誌、教科書、SEC提出書類、金融NLPタスク、投資の質問から1,335件の指示データセットを構築する。
  • InvestLMをGPT-3.5、GPT-4、Claude-2、オープンベースラインと、専門家評価とGPT-4スタイルのスコアリングを通じて比較する。
Figure 1: Expert evaluation.
Figure 1: Expert evaluation.

実験結果

リサーチクエスチョン

  • RQ1小規模で厳選された金融ドメイン指示セットは、強力な基盤モデルを高品質な金融LLMへと変換できるか。
  • RQ2ドメイン指示チューニングは、金融NLPタスクにおける汎用指示データと比較して性能にどのような影響を与えるか。
  • RQ3InvestLMは金融ベンチマークと専門家評価において、最先端の商用モデルと比較してどの程度のパフォーマンスを示すか。
  • RQ4チューニングで明示的に使用されていない金融NLPタスクへの一般化性はどの程度か。
  • RQ5ドメイン調整済みLLaMAベースモデルとベースラインLLaMAの投資シナリオにおける挙動の違いは何か。

主な発見

  • InvestLMの専門家評価済み回答は、テスト質問の結果でGPT-3.5およびGPT-4と同等、またはそれを上回る。
  • InvestLMは、チューニングで使用されなかったいくつかの金融NLPベンチマークで強い一般化を示す。
  • ドメイン指示チューニングは、より小さなモデル(7B)でより大きなモデル(65B)よりも大きなゲインをもたらす。
  • 汎用的なAlpaca風指示は、ドメインタスクの性能を損なう可能性があり、ドメイン特異データの価値を強調する。
  • InvestLMはLLaMAに比べ幻覚を抑制し、簡潔で論理的な投資結論を提供する。
  • InvestLMの結果は、ターゲット指示による高ドメイン性能を示す「表面的適合仮説」と整合する。
Figure 2: GPT-4 evaluation.
Figure 2: GPT-4 evaluation.

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。