Skip to main content
QUICK REVIEW

[論文レビュー] Qibo: A Large Language Model for Traditional Chinese Medicine

Heyi Zhang, Xin Wang|arXiv (Cornell University)|Mar 24, 2024
Traditional Chinese Medicine Studies被引用数 13
ひとこと要約

Qiboは、事前学習から監視付きファインチューニングまでをカスタムのTraditional Chinese Medicineコーパスで訓練した、Chinese-LLaMAベースのLLMであり、TCMの理解と応用を評価する専用の評価ベンチマークを備えています。

ABSTRACT

Large Language Models (LLMs) has made significant progress in a number of professional fields, including medicine, law, and finance. However, in traditional Chinese medicine (TCM), there are challenges such as the essential differences between theory and modern medicine, the lack of specialized corpus resources, and the fact that relying only on supervised fine-tuning may lead to overconfident predictions. To address these challenges, we propose a two-stage training approach that combines continuous pre-training and supervised fine-tuning. A notable contribution of our study is the processing of a 2GB corpus dedicated to TCM, constructing pre-training and instruction fine-tuning datasets for TCM, respectively. In addition, we have developed Qibo-Benchmark, a tool that evaluates the performance of LLM in the TCM on multiple dimensions, including subjective, objective, and three TCM NLP tasks. The medical LLM trained with our pipeline, named $ extbf{Qibo}$, exhibits significant performance boosts. Compared to the baselines, the average subjective win rate is 63%, the average objective accuracy improved by 23% to 58%, and the Rouge-L scores for the three TCM NLP tasks are 0.72, 0.61, and 0.55. Finally, we propose a pipline to apply Qibo to TCM consultation and demonstrate the model performance through the case study.

研究の動機と目的

  • Traditional Chinese Medicine (TCM) の理論と現代のAIモデルとのギャップに対処するため、TCMに特化したLLMの創出を動機づける。
  • TCMに焦点を当てたLLMのための、pre-trainingからSFTまでの完全なトレーニングパイプラインを開発する。
  • より良い領域知識の獲得のために、高品質で多様なTCMコーパスとデータ処理ルールを構築する。
  • TCMに関する知識・推論・安全性を定量的に評価するための Qibo-benchmark を作成する。

提案手法

  • 多様なTCMおよび現代医療コーパスで継続的な事前学習を行い、Chinese-LLaMA に領域知識と診断推論を組み込む。
  • Alpaca風の対話形式に翻訳された複数ソースのTCM対話データセットを用いた監視付きファインチューニング (SFT)。
  • SFT のための4つのデータソースの組み込み:TCMの単発対話と多 turns対話、TCM NLPタスクデータ、および一般的な医療対話。領域適合性と一般化を確保。
  • 高品質なトレーニングコーパスを作成するための、文字レベルと段落レベルの多層クレンジングと手動検証を備えたデータ処理パイプライン。
  • GPT-4と人間の専門家評価による安全性、専門性、流暢さの評価に加え、客観的なNLPタスクベンチマークを活用。

実験結果

リサーチクエスチョン

  • RQ1How effectively can a LLaMA-based model be trained from pre-training to SFT to specialize in Traditional Chinese Medicine?
  • RQ2Does a domain-specific corpus and evaluation benchmark improve the model's TCM understanding, dialectics, and prescription recognition compared to general medical LLMs?
  • RQ3Can Qibo outperform existing Chinese medical LLMs in TCM-specific tasks and multi-turn dialogues while maintaining safety and professionalism?
  • RQ4What are the limitations and safety considerations when deploying a TCM-focused LLM in practice?

主な発見

  • Qibo demonstrates strong domain performance in traditional Chinese medicine across subjective (professionalism, safety, fluency) and objective assessments.
  • Qibo-13B yields higher accuracy than Qibo-7B on multi-turn TCM evaluation tasks, indicating scale benefits within the framework.
  • In TCM NLP tasks, Qibo often surpasses baseline medical models but may lag task-specific models optimized for those tasks.
  • The authors provide a dedicated Qibo-benchmark enabling quantitative comparison of TCM understanding and application across models.
  • The model acknowledges limitations, including potential inaccuracies and the need for cautious real-world medical use, and points to future safety and multi-modal integration work.

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。