Skip to main content
QUICK REVIEW

[論文レビュー] Fine-tuning large language models for domain adaptation: Exploration of training strategies, scaling, model merging and synergistic capabilities

Wei Lu, Rachel K. Luu|arXiv (Cornell University)|Sep 5, 2024
Natural Language Processing Techniques被引用数 4
ひとこと要約

本論文は、大規模言語モデル(LLM)を材料科学および工学分野に適応させるためのファインチューニング戦略—継続的事前学習(CPT)、教師ありファインチューニング(SFT)、および好み最適化(DPO、ORPO)—を調査している。モデルメルジングが個々のモデルを超える非線形的で相乗的な能力を生み出すことを示しており、Llama 3.1 8BおよびMistral 7Bでは性能向上が観察された。一方、小型モデル(1.7B)では限られた顕在化を示しており、このような相乗効果にはモデルスケーリングが不可欠であることが示唆される。

ABSTRACT

The advancement of Large Language Models (LLMs) for domain applications in fields such as materials science and engineering depends on the development of fine-tuning strategies that adapt models for specialized, technical capabilities. In this work, we explore the effects of Continued Pretraining (CPT), Supervised Fine-Tuning (SFT), and various preference-based optimization approaches, including Direct Preference Optimization (DPO) and Odds Ratio Preference Optimization (ORPO), on fine-tuned LLM performance. Our analysis shows how these strategies influence model outcomes and reveals that the merging of multiple fine-tuned models can lead to the emergence of capabilities that surpass the individual contributions of the parent models. We find that model merging leads to new functionalities that neither parent model could achieve alone, leading to improved performance in domain-specific assessments. Experiments with different model architectures are presented, including Llama 3.1 8B and Mistral 7B models, where similar behaviors are observed. Exploring whether the results hold also for much smaller models, we use a tiny LLM with 1.7 billion parameters and show that very small LLMs do not necessarily feature emergent capabilities under model merging, suggesting that model scaling may be a key component. In open-ended yet consistent chat conversations between a human and AI models, our assessment reveals detailed insights into how different model variants perform and show that the smallest model achieves a high intelligence score across key criteria including reasoning depth, creativity, clarity, and quantitative precision. Other experiments include the development of image generation prompts based on disparate biological material design concepts, to create new microstructures, architectural concepts, and urban design based on biological materials-inspired construction principles.

研究の動機と目的

  • CPT、SFT、および好みベース最適化(DPO、ORPO)がドメイン特化の科学的応用におけるLLM性能に与える影響を評価すること。
  • モデルメルジングが個々の親モデルに存在しなかった顕在的能力を生み出すかどうかを調査すること。
  • 親モデル間のスケーリングと多様性が、相乗的性能向上を可能にする役割を評価すること。
  • オープンエンドでマルチターンの会話および画像生成タスクにおいて、メルジッドモデルの有効性を評価すること。生物的インスピレーションを受けるプロンプトを用いる。
  • 1.7Bパラメータの小型LLMがメルジングによって顕在的能力を示す可能性を検討し、大規模モデルと比較すること。

提案手法

  • 材料科学および工学分野のドメイン特化した科学的コーパスを用いて、Llama 3.1 8BおよびMistral 7BをCPTで訓練した。
  • キュレートされたインstructデータセットを用いてSFTを実施し、タスク固有の性能およびインstructフォローアップ能力を向上させた。
  • DPOおよびORPOを用いて好みベース最適化を実施し、複雑な報酬モデリングを回避しながら好み信号を直接最適化した。
  • 効率的なファインチューニングのためLoRAを採用し、すべての線形層にランクr = 16または64の低ランクアダプタを適用した。
  • パラメータ重み付き平均化を用いてファインチューニング済みモデルをメルジングし、標準化スコア(ZAi)およびクラスタリングを用いて性能を分析した。
  • マルチターンのヒューマン-AI会話および生物的インスピレーションを受けるプロンプトを用いた、ファインチューニング済みのFLUX.1-devモデルを用いた画像生成を評価した。

実験結果

リサーチクエスチョン

  • RQ1CPT、SFT、および好みベース最適化(DPO、ORPO)は、材料科学分野のLLM性能向上において、どのように比較されるか?
  • RQ2複数のファインチューニング済みモデルをメルジングすることで、個々の親モデルに存在しなかった顕在的能力が生じるか?
  • RQ3モデルメルジングにおける顕在的相乗的行動を可能にするために、モデルスケーリングはどのような役割を果たすか?
  • RQ4親モデル間の多様性は、モデルメルジングの成功にどのように影響するか?
  • RQ5非常に小型のLLM(例:1.7Bパラメータ)は、メルジング後に顕在的能力を示すか、それともスケーリングが前提条件となるか?

主な発見

  • モデルメルジングにより、親モデルの平均を上回る性能向上が見られ、モデルパラメータ間の非線形的で相乗的な相互作用が示された。
  • 最良の性能を示したメルジッドモデルは、期待される平均を大幅に上回る標準化実スコア(ZAi)を達成しており、顕在的能力の存在を示している。
  • 大規模モデル(Llama 3.1 8B、Mistral 7B)は、メルジングによって強い相乗的性能を示したが、1.7Bモデルは顕在的行動が限定的であった。
  • 最小のモデル(1.7B)は、推論、創造性、明確さ、正確性の面でヒューマン-AI会話において高い知能スコアを達成したが、顕在的能力は限られていた。
  • CPTおよびSFTと組み合わせた好みベース最適化(ORPO、DPO)は、SFT単体よりも優れた性能を示し、特にインstructフォローアップおよび推論能力で顕著であった。
  • 生物的インスピレーションを受けるプロンプトを用いたメルジッドモデルによる画像生成は、抽象的で有機的なマイクロ構造を効果的に生成し、マルチモodal推論能力を示している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。