Skip to main content
QUICK REVIEW

[論文レビュー] New Methods for Metadata Extraction from Scientific Literature

Dominika Tkaczyk|arXiv (Cornell University)|Oct 27, 2017
Handwritten Text Recognition Techniques参考文献 4被引用数 8
ひとこと要約

本稿では、生デジタル科学論文から自動的で正確かつ柔軟なメタデータ抽出を実現する教師ありおよび教師なし機械学習ベースのアルゴリズムを提案する。文書レイアウト解析、領域分類、メタデータ、参考文献、本文の構造的解析を統合することで、多様なレイアウトにおいて高い正確性を達成し、評価において既存手法を上回る性能を示した。

ABSTRACT

Within the past few decades we have witnessed digital revolution, which moved scholarly communication to electronic media and also resulted in a substantial increase in its volume. Nowadays keeping track with the latest scientific achievements poses a major challenge for the researchers. Scientific information overload is a severe problem that slows down scholarly communication and knowledge propagation across the academia. Modern research infrastructures facilitate studying scientific literature by providing intelligent search tools, proposing similar and related documents, visualizing citation and author networks, assessing the quality and impact of the articles, and so on. In order to provide such high quality services the system requires the access not only to the text content of stored documents, but also to their machine-readable metadata. Since in practice good quality metadata is not always available, there is a strong demand for a reliable automatic method of extracting machine-readable metadata directly from source documents. This research addresses these problems by proposing an automatic, accurate and flexible algorithm for extracting wide range of metadata directly from scientific articles in born-digital form. Extracted information includes basic document metadata, structured full text and bibliography section. Designed as a universal solution, proposed algorithm is able to handle a vast variety of publication layouts with high precision and thus is well-suited for analyzing heterogeneous document collections. This was achieved by employing supervised and unsupervised machine-learning algorithms trained on large, diverse datasets. The evaluation we conducted showed good performance of proposed metadata extraction algorithm. The comparison with other similar solutions also proved our algorithm performs better than competition for most metadata types.

研究の動機と目的

  • 増加する学術文献の処理を効率化することで、科学的情報過多の課題に対処する。
  • デジタル科学的出版物における高品質で機械可読なメタデータの不足を克服する。
  • 多様な文書レイアウトを高精度に処理できる汎用的で強固なソリューションを開発する。
  • 現代の研究インfraストラクチャが、引用ネットワーク、類似推薦、インパacts分析などの知的サービスを提供できるようにする。
  • 生デジタル論文から構造的メタデータ、参考文献、本文をスケーラブルかつ正確に抽出するパイプラインを提供する。

提案手法

  • 教師ありおよび教師なし機械学習を組み合わせたハイブリッドアプローチを採用し、文書レイアウト解析と領域分類を実施する。
  • 空間的近接性と最近傍距離解析に基づき、読み順の特定にドクストラムアルゴリズムを用いる。
  • 幾何学的特徴、テキスト的特徴、文脈的特徴を用いて、領域分類モデル(例:タイトル、著者、要旨、参考文献)を階層的に適用する。
  • 段階的なパースパイプラインを実装する:ページセグメンテーション → コンテンツ分類 → メタデータ抽出 → 参考文献および本文構造の解析。
  • GROTOAP2 などの大規模で多様なトレーニングデータセットを活用し、機関情報および引用情報のパースに高い正確性を実現する。
  • モジュール型のコンポONENTをメタデータ、参考文献、本文抽出に統合し、柔軟性と拡張性を実現する。

実験結果

リサーチクエスチョン

  • RQ1生デジタルフォーマットの多様な科学的論文レイアウトから、どのように高精度なメタデータ抽出を実現できるか?
  • RQ2どのような機械学習技術が、多様な出版スタイルにわたる文書領域の分類を強固に可能にするか?
  • RQ3統合されたシステムは、複数のメタデータタイプにおいて、既存のメタデータ抽出ツールと比較して優れた性能を発揮できるか?
  • RQ4提案されたレイアウト解析および読み順特定の有効性は、文書構造の保持にどの程度効果的か?
  • RQ5このシステムは、異なる科学分野および文書タイプにどの程度一般化可能か?

主な発見

  • 提案手法は、PMCおよびElsevierデータセットの両方において、基本的メタデータ、著者情報、参考文献抽出の分野で、既存手法を上回る性能を示した。
  • GROTOAP2 データセットでは、コンテンツ分類で F スコア 0.92、メタデータ分類で F スコア 0.89 を達成した。
  • 参考文献パーサーは GROTOAP2 引用データセットで F スコア 0.87 を達成し、構造的参照抽出の高精度を示した。
  • 処理遅延はページあたり平均 1.2 秒であり、処理時間の 70% がコンテンツ分類およびレイアウト解析に費やされた。
  • 分類タスクの混同行列は、特にタイトル、要旨、参考文献セクションで低い誤差率を示した。
  • 評価により、本手法が多様なレイアウトおよび出版タイプにわたり堅牢であり、複数のテストコレクションで一貫した性能を発揮することが確認された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。