Skip to main content
QUICK REVIEW

[論文レビュー] Measuring Patent Claim Generation by Span Relevancy

Jieh-Sheng Lee, Jieh Hsiang|arXiv (Cornell University)|Aug 26, 2019
Topic Modeling参考文献 6被引用数 5
ひとこと要約

本稿では、スパンの関連性を用いて特許出願文書生成の質を定量的に測定する、新しいスパンベースのフレームワークを提案する。特許出願の関連性を自然言語推論問題として扱い、BERTを微調整してスパンペアの関連性を分類し、GPT-2を微調整して出願文書を生成することで、生成の多様性が高まるにつれてスパン関連性が低下することを実証した。これにより、AI支援特許作成における「拡張的発明(augmented inventing)」のための信頼できる定量的指標であることが裏付けられた。

ABSTRACT

Our goal of patent claim generation is to realize "augmented inventing" for inventors by leveraging latest Deep Learning techniques. We envision the possibility of building an "auto-complete" function for inventors to conceive better inventions in the era of artificial intelligence. In order to generate patent claims with good quality, a fundamental question is how to measure it. We tackle the problem from a perspective of claim span relevancy. Patent claim language was rarely explored in the NLP field. It is unique in its own way and contains rich explicit and implicit human annotations. In this work, we propose a span-based approach and a generic framework to measure patent claim generation quantitatively. In order to study the effectiveness of patent claim generation, we define a metric to measure whether two consecutive spans in a generated patent claims are relevant. We treat such relevancy measurement as a span-pair classification problem, following the concept of natural language inference. Technically, the span-pair classifier is implemented by fine-tuning a pre-trained language model. The patent claim generation is implemented by fine-tuning the other pre-trained model. Specifically, we fine-tune a pre-trained Google BERT model to measure the patent claim spans generated by a fine-tuned OpenAI GPT-2 model. In this way, we re-use two of the state-of-the-art pre-trained models in the NLP field. Our result shows the effectiveness of the span-pair classifier after fine-tuning the pre-trained model. It further validates the quantitative metric of span relevancy in patent claim generation. Particularly, we found that the span relevancy ratio measured by BERT becomes lower when the diversity in GPT-2 text generation becomes higher.

研究の動機と目的

  • 自然言語処理分野における特許出願文書生成の質を評価する定量的指標の不足に対処すること。
  • AI支援の自動補完システムを活用して「拡張的発明(augmented inventing)」を可能にすること。
  • 自然言語推論の原則を用いて特許出願文言をスパンペアの関連性タスクとしてモデル化すること。
  • 事前学習済み言語モデルを用いて、出願文書生成の質を測る新しい指標を検証すること。
  • 生成の多様性と特許文書における出願文の整合性の関係を分析すること。

提案手法

  • 特許出願文書生成の評価を、連続するスパン間の関連性をモデル化するスパンペア分類タスクとして定式化する。
  • 特許出願文書内の連続するスパンペアの意味的関連性を分類するために、事前学習済みBERTモデルを微調整する。
  • 別個に微調整されたOpenAI GPT-2モデルを用いて、評価用の特許出願文書を生成する。
  • 実際の特許出願文書から抽出したラベル付きスパンペアを用いて、スパンペア分類器を学習させ、出願文の整合性を模擬する。
  • 最先端の事前学習済みモデル(BERTとGPT-2)を用いて、出願文書生成および関連性スコアリングの両方で転移学習を適用する。
  • スパン関連性比を出願文書の質の定量的指標として測定し、生成された出願文書全体における整合性を反映する。

実験結果

リサーチクエスチョン

  • RQ1スパンレベルの関連性を用いて、どのように特許出願文書生成の質を定量的に測定できるか?
  • RQ2GPT-2によるテキスト生成の多様性が高まると、生成された特許出願文書の整合性にどの程度影響を与えるか?
  • RQ3微調整されたBERTモデルは、特許出願文書のスパン関連性を効果的に分類できるか?
  • RQ4スパン関連性は、特許出願文書生成タスクにおける全体的な出願文書の質の信頼できる代理指標とみなせるか?
  • RQ5特許出願文書の構造的特徴(明示的・暗黙的の注釈を含む)は、なぜスパンベースの評価を支援するのか?

主な発見

  • 微調整されたBERTモデルは、特許出願文書におけるスパン関連性の分類を効果的に学習し、提案された指標の実現可能性を示した。
  • GPT-2で生成された出願文書の多様性が高まるにつれて、スパン関連性比が低下することが判明し、創造性と整合性のトレードオフが示された。
  • 提案されたスパンペア分類フレームワークは、特許出願文書生成の質を評価する信頼できる定量的指標を提供する。
  • 特許出願文書の言語には、事前学習済み言語モデルを用いたスパンベースの評価を可能にする十分な構造的・意味的手がかりが含まれている。
  • 結果から、スパン関連性がAI生成特許出願文書の整合性と質を評価する意味のある指標であることが裏付けられた。
  • このフレームワークにより、特許出願文書生成の客観的・自動的評価が可能となり、AI支援発明ツールの開発を支援する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。