Skip to main content
QUICK REVIEW

[論文レビュー] Symbolic Knowledge Distillation: from General Language Models to Commonsense Models

Peter West, Chandra Bhagavatula|arXiv (Cornell University)|Oct 14, 2021
Topic Modeling被引用数 7
ひとこと要約

この論文では、GPT-3を用いて人為的知識なしに高品質な常識知識グラフ(Atomic${}^{\text{10x}}$)を自動生成し、コンactなニューラル常識モデル(Comet${}^{\text{dis}}_{\text{til}}$)を訓練するためのフレームワーク、シンボリック知識蒸留(Symbolic Knowledge Distillation)を提案する。100倍も小さいにもかかわらず、学生モデルは人為的アノテーションによるAtomic${}^{20}$コーパスおよびその教師モデルであるGPT-3を、スケール、品質、多様性の面で上回る常識推論性能を示す。

ABSTRACT

The common practice for training commonsense models has gone from-human-to-corpus-to-machine: humans author commonsense knowledge graphs in order to train commonsense models. In this work, we investigate an alternative, from-machine-to-corpus-to-machine: general language models author these commonsense knowledge graphs to train commonsense models. Our study leads to a new framework, Symbolic Knowledge Distillation. As with prior art in Knowledge Distillation (Hinton et al., 2015), our approach uses larger models to teach smaller models. A key difference is that we distill knowledge symbolically-as text-in addition to the neural model. We also distill only one aspect-the commonsense of a general language model teacher, allowing the student to be a different type, a commonsense model. Altogether, we show that careful prompt engineering and a separately trained critic model allow us to selectively distill high-quality causal commonsense from GPT-3, a general language model. Empirical results demonstrate that, for the first time, a human-authored commonsense knowledge graph is surpassed by our automatically distilled variant in all three criteria: quantity, quality, and diversity. In addition, it results in a neural commonsense model that surpasses the teacher model's commonsense capabilities despite its 100x smaller size. We apply this to the ATOMIC resource, and share our new symbolic knowledge graph and commonsense models.

研究の動機と目的

  • 人為的常識知識グラフに依存しない、マシンからコーパスへ、再びマシンへとつながるパイプラインの開発。
  • 大規模言語モデルを教師として活用することで、常識知識の品質とスケーラビリティを向上させること。
  • 一般言語モデルから、特定の専用のコンパクトなニューラルモデルに、因果的常識のみを蒸留することの実現。
  • 評価モデルを用いて、自動生成されたシンボリック知識の品質を評価・向上させること。
  • 機械生成された知識が、スケール、多様性、正確性の面で人為的構築知識を凌駕できることを実証すること。

提案手法

  • 因果的推論に焦点を当てた巧みなプロンプトを設計し、GPT-3を教師モデルとして用いて、シンボリックな常識トリプルを生成する。
  • GPT-3の出力テキストから、人間が読みやすく解釈可能なシンボリック知識グラフ、Atomic${}^{\text{10x}}$を構築する。
  • 生成されたトリプルの品質を評価するための別個の評価モデル(クリティックモデル)を訓練し、高品質な知識の選択的蒸留を可能にする。
  • ニューラル学生モデル(Comet${}^{\text{dis}}_{\text{til}}$)への知識蒸留に加え、シンボリック知識グラフへの蒸留も実施し、二重の転送を可能にする。
  • 人為ラベルデータを用いて評価モデルをファインチューニングし、品質フィルタリングの信頼性と正確性を向上させる。
  • 蒸留されたシンボリックコーパス上で、コンパクトな常識モデル(Comet${}^{\text{dis}}_{\text{til}}$)を訓練し、計算負荷を低減しつつ高い精度での推論を実現する。

実験結果

リサーチクエスチョン

  • RQ1GPT-3のような一般言語モデルが、因果的常識知識をシンボリック知識グラフに蒸留する高品質な教師として機能できるか?
  • RQ2機械生成知識で訓練されたニューラル学生モデルが、人為的アノテーションデータで訓練されたモデルを上回れるか?
  • RQ3評価モデルの導入により、直接生成する場合と比較して、蒸留されたシンボリック知識の品質が顕著に向上するか?
  • RQ4機械生成された常識知識が、スケール、多様性、正確性の面で人為的知識を凌駆できるか?
  • RQ5一般言語モデルから、特定の側面(因果的常識)のみを、異なるタイプのモデル(常識モデル)に蒸留することが可能か?

主な発見

  • 機械生成された知識グラフ、Atomic${}^{\text{10x}}$は、7種類の常識推論タイプにわたり、数量、品質、多様性のすべての基準で、人為的アノテーションによるAtomic${}^{20}$を上回る。
  • GPT-3の100分の1のサイズである学生モデルComet${}^{\text{dis}}_{\text{til}}$は、教師モデルGPT-3を上回る常識推論の正確性を達成している。
  • 評価モデルは、蒸留知識の品質を顕著に向上させ、人為的アノテーションの基準を凌駆する高精度なシンボリックトリプルの生成を可能にした。
  • 100件の生成例を手動で検査した結果、成人向けコンテンツが1件のみ検出されたことから、限定的で制御されたプロンプトにより、有害または不適切な出力のリスクが極めて低いことが示された。
  • 本フレームワークは、人間が作成者ではなく評価者として参加する協働的で人間と機械の連携パイプラインを可能にし、コスト削減とスケーラビリティの向上を実現する。
  • シンボリック知識蒸留フレームワークは、常識推論タスクにおいて、機械生成知識が人為的構築知識よりも包括的かつ正確である可能性を示している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。