Skip to main content
QUICK REVIEW

[論文レビュー] PIXIU: A Large Language Model, Instruction Data and Evaluation Benchmark for Finance

Qianqian Xie, Weiguang Han|arXiv (Cornell University)|Jun 8, 2023
Stock Market Forecasting Methods被引用数 43
ひとこと要約

PIXIUは、LLaMAからファインチューニングした大規模マルチタスク・マルチモーダル指示データセット(FIT)と評価ベンチマーク(FLARE)を備えたオープンソースの金融LLM FinMAを提示する。FinMAは金融NLPタスクで優れているが、複雑なQAと株価動向予測には改善の余地がある。

ABSTRACT

Although large language models (LLMs) has shown great performance on natural language processing (NLP) in the financial domain, there are no publicly available financial tailtored LLMs, instruction tuning datasets, and evaluation benchmarks, which is critical for continually pushing forward the open-source development of financial artificial intelligence (AI). This paper introduces PIXIU, a comprehensive framework including the first financial LLM based on fine-tuning LLaMA with instruction data, the first instruction data with 136K data samples to support the fine-tuning, and an evaluation benchmark with 5 tasks and 9 datasets. We first construct the large-scale multi-task instruction data considering a variety of financial tasks, financial document types, and financial data modalities. We then propose a financial LLM called FinMA by fine-tuning LLaMA with the constructed dataset to be able to follow instructions for various financial tasks. To support the evaluation of financial LLMs, we propose a standardized benchmark that covers a set of critical financial tasks, including five financial NLP tasks and one financial prediction task. With this benchmark, we conduct a detailed analysis of FinMA and several existing LLMs, uncovering their strengths and weaknesses in handling critical financial tasks. The model, datasets, benchmark, and experimental results are open-sourced to facilitate future research in financial AI.

研究の動機と目的

  • オープンで指示に従う金融LLMと厳選されたデータセットの必要性を動機づける。
  • マルチタスク・マルチモーダルな金融指示でLLaMAをファインチューニングしてFinMAを作成する。
  • 最初の大規模金融指示調整データセット(FIT)を開発する(136Kサンプル)。
  • 金融NLPと予測タスクを含む包括的なベンチマークFLAREを提案する。
  • 金融AI研究を推進するためのオープンリソースを提供する。

提案手法

  • NLPと株価予測を網羅するオープンソースの金融データセットから、テキスト、表、時系列などのマルチモーダルデータを含むFITを構築する。
  • タスクごとにドメイン特化の指示を設計し、指示-テキスト/文脈-応答という構成の指示チューニングサンプルを組み立てる。
  • AdamW、指定されたハイパーパラメータと多エポックスケジュールを用いて FIT 上で LLaMA モデル(7B および 30B、さらに 7B-full バリアント)をファインチューニングする。
  • FLARE を、6データセットにまたがる4つの金融NLPタスクと、3データセットにまたがる1つの金融予測タスクで作成する。タスクごとに標準評価指標。
  • FLARE におけるゼロショットおよびFew-shot設定で、FinMAをBloombergGPT、GPT-4、ChatGPT、BLOOM、GPT-NeoX、OPT-66B、Vicuna-13Bと比較する。

実験結果

リサーチクエスチョン

  • RQ1オープンな金融指示データとモデルは、金融分野のプロプライエタリLLMとのギャップを埋めることができるか?
  • RQ2マルチタスク・マルチモーダルな指示チューニングはFinMAの金融タスクの性能にどう影響するか?
  • RQ3金融NLPおよび予測タスクにおけるFinMAの強みと限界は、ベースラインと比較してどうか?
  • RQ4モデルサイズと指示データの品質はFLAREタスク全体の性能にどのような影響を与えるか?

主な発見

  • FinMAはFPB、FiQA-SA、およびHeadline NLPタスクで他のいくつかのLLMを顕著に上回る(例:FinMA-30BはFPBでGPT-4を約10%のF1で上回り、BloombergGPTを約37%のF1上回る)。
  • FinMAはNERタスクでも競争力のある結果を達成し、BloombergGPT等をNERタスクで上回る。
  • 複雑な数値推論タスク(FinQA, ConvFinQA)では、LLaMAというバックボーンモデルの定量的推論の限界のため、FinMAはGPT-4やBloombergGPTに遅れをとる。
  • 株価動向予測ではすべてのLLMが限られた性能を示す。FinMA-7B-fullはACL18で改善を示すが、他のデータセットでは依然として弱く、タスクの難しさを浮き彫りにしている。
  • FinMA-full(NLPおよび予測データで訓練)はACL18で最も高い性能を示し、競争力のあるNLP結果を示しており、ドメイン適合・タスク総合的なファインチューニングの価値を示している。
  • 本研究は、多くのタスクにおいて指示データの品質とタスクの整合性が、単にモデルサイズを増やすよりも重要である可能性を強調している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。