Skip to main content
QUICK REVIEW

[논문 리뷰] GenericsKB: A Knowledge Base of Generic Statements

Sumithra Bhakthavatsalam, Chloe Anastasiades|arXiv (Cornell University)|2020. 05. 02.
Natural Language Processing Techniques참고 문헌 19인용 수 44
한 줄 요약

GenericsKB는 주제 메타데이터와 학습된 신뢰도가 포함된 자연 발생 일반 진술의 대규모 컬렉션으로, GenericsKB-Best(1M+ 문장)는 합성 일반 진술로 보강되었다; 이는 QA 설명을 개선하고 더 큰 일반 코퍼스를 사용할 때보다 추론 과제에서 성능을 높일 수 있다(3.5M+ 포함).

ABSTRACT

We present a new resource for the NLP community, namely a large (3.5M+ sentence) knowledge base of *generic statements*, e.g., "Trees remove carbon dioxide from the atmosphere", collected from multiple corpora. This is the first large resource to contain *naturally occurring* generic sentences, as opposed to extracted or crowdsourced triples, and thus is rich in high-quality, general, semantically complete statements. All GenericsKB sentences are annotated with their topical term, surrounding context (sentences), and a (learned) confidence. We also release GenericsKB-Best (1M+ sentences), containing the best-quality generics in GenericsKB augmented with selected, synthesized generics from WordNet and ConceptNet. In tests on two existing datasets requiring multihop reasoning (OBQA and QASC), we find using GenericsKB can result in higher scores and better explanations than using a much larger corpus. This demonstrates that GenericsKB can be a useful resource for NLP applications, as well as providing data for linguistic studies of generics and their semantics. GenericsKB is available at https://allenai.org/data/genericskb.

연구 동기 및 목표

  • 자연적으로 발생하는 대규모 일반 진술 코퍼스를 NLP 및 언어학용으로 제공합니다.
  • 주제 메타데이터, 주변 맥락, 학습된 신뢰도 점수로 진술을 주석합니다.
  • WordNet 및 ConceptNet에서 보강된 합성 일반 진술을 포함한 고품질 서브셋(GenericsKB-Best)을 공개합니다.
  • 질문 응답 및 설명 생성과 같은 다운스트림 작업에서 GenericsKB의 활용성을 입증합니다.

제안 방법

  • 세 코퍼스(Waterloo, SimpleWikipedia, ARC)에서 총 1.7B 문장으로부터 문장을 수집합니다.
  • 정규식(regex), 길이 휴리스틱, 언어 탐지를 사용하여 노이즈를 정리하고 필터링합니다.
  • 27개의 수작업으로 작성된 어휘-통사 규칙으로 독립적인 일반 진술을 식별합니다.
  • 일반 진실로서의 유용성에 대한 크라우드 주석 판단으로 학습된 BERT 분류기로 후보 일반 진술에 점수를 부여합니다.

실험 결과

연구 질문

  • RQ1크고 자연 발생적인 일반 진술 코퍼스를 높은 품질과 맥락적 완전성을 갖고 구축할 수 있는가?
  • RQ2GenericsKB가 더 큰 일반 코퍼스에 비해 기존 다중 호의 추론 태스크에서 성능이나 설명을 개선하는가?
  • RQ3QA 및 설명 작업에서 GenericsKB의 품질과 활용도는 어떠한가?

주요 결과

  • 최종 GenericsKB는 맥락 메타데이터, 맥락, 신뢰도 점수가 포함된 3,433,000문장을 담고 있다.
  • GenericsKB-Best는 WordNet/ConceptNet 데이터로 보강된 1,020,868개의 일반 진술(GenericsKB에서 774,621개, 합성 246,247개)을 포함한다.
  • OpenBookQA에서 GenericsKB-Best는 GenericsKB(0.632) 및 QASC-17M baseline(0.660)보다 더 높은 QA 성능(0.678)을 보인다.
  • GenericsKB-Best는 QASC에서 두 호 설명을 훨씬 더 잘 생성한다(0.61 대 QASC-17M의 0.44; 다른 지표에서 0.79 대 0.66).
  • 주석/품질 검사는 GenericsKB-Best 샘플의 유용성 기준에 대해 약 85%의 일치를 보였으며, 맥락 누출이나 모호성은 상대적으로 적은 편임을 시사한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.