Skip to main content
QUICK REVIEW

[논문 리뷰] On the comparability of Pre-trained Language Models

Matthias Aßenmacher, Christian Heumann|arXiv (Cornell University)|2020. 01. 03.
Topic Modeling참고 문헌 41인용 수 11
한 줄 요약

이 논문은 2018–2019년도의 최신 사전 훈련된 언어 모델들을 종합적으로 비교 분석하며, 아키텍처 혁신, 훈련 자원, 전이 학습 통합에 중점을 둡니다. 연구는 재현 가능성과 공정한 벤치마킹을 향상시키기 위해 표준화된 보고 방식을 제안하며, 모델 개발 및 평가 지표의 투명성 향상을 위한 실현 가능하고 실행 가능한 지침을 제시합니다.

ABSTRACT

Recent developments in unsupervised representation learning have successfully established the concept of transfer learning in NLP. Mainly three forces are driving the improvements in this area of research: More elaborated architectures are making better use of contextual information. Instead of simply plugging in static pre-trained representations, these are learned based on surrounding context in end-to-end trainable models with more intelligently designed language modelling objectives. Along with this, larger corpora are used as resources for pre-training large language models in a self-supervised fashion which are afterwards fine-tuned on supervised tasks. Advances in parallel computing as well as in cloud computing, made it possible to train these models with growing capacities in the same or even in shorter time than previously established models. These three developments agglomerate in new state-of-the-art (SOTA) results being revealed in a higher and higher frequency. It is not always obvious where these improvements originate from, as it is not possible to completely disentangle the contributions of the three driving forces. We set ourselves to providing a clear and concise overview on several large pre-trained language models, which achieved SOTA results in the last two years, with respect to their use of new architectures and resources. We want to clarify for the reader where the differences between the models are and we furthermore attempt to gain some insight into the single contributions of lexical/computational improvements as well as of architectural changes. We explicitly do not intend to quantify these contributions, but rather see our work as an overview in order to identify potential starting points for benchmark comparisons. Furthermore, we tentatively want to point at potential possibilities for improvement in the field of open-sourcing and reproducible research.

연구 동기 및 목표

  • 아키텍처, 훈련 데이터, 전이 학습 통합 측면에서 주요 사전 훈련된 언어 모델들(예: BERT, GPT, RoBERTa) 간의 핵심 차이를 명확히 하기.
  • 성능 향상에 기여하는 아키텍처 변경, 모델 크기, 계산 자원의 상대적 기여도를 정량화하지 않고도 파악하기.
  • 특히 계산 비용, 하이퍼파rameter 튜닝, 자원 사용에 관해 일관된 보고 표준이 부족한 NLP 연구의 문제를 해결하기.
  • 표준화된 보고 체계를 통해 재현 가능성과 개방 과학을 촉진하기 위해 NLP 커뮤니티 전반에 걸쳐 실용적인 프레임워크를 제안하기.
  • 광범위한 하이퍼파rameter 튜닝과 대규모 모델 훈련의 환경적 및 실용적 영향을 부각시키며, 개발 자원 보고의 투명성 강화를 촉구하기.

제안 방법

  • 구조화된 표를 활용해 아키텍처 및 훈련 특성에 기반해 12개의 사전 훈련된 언어 모델(예: Word2Vec, FastText, ULMFiT, ELMo, GPT, BERT, GPT2, XLNet, RoBERTa, ALBERT)을 체계적으로 비교.
  • 컨텍스트성(일방향 대 양방향), 훈련 파라다임(단일 임bedding 전용 대 종합적 훈련 가능), 아키텍처(LSTM, Transformer 등) 기반으로 모델을 분류.
  • 각 보고 차원(예: 모델 아키텍처, 파라미터 수, 하이퍼파ram터 등)에 대해 실현 가능성, 현재 구현 수준, 재현 가능성에 대한 관련성을 평가하는 보고 프레임워크를 제안.
  • 환경적 및 실용적 고려사항을 뒷받침하기 위해 어휘 자원(예: 훈련 코퍼스 크기, 가용성)과 계산 비용(GPU 시간, 에너지 소비) 보고의 중요성을 강조.
  • 정확한 FLOP 계산 대신, 훈련 노력 비교에 실용적이고 접근 가능한 대안으로 OpenAI 블로그의 GPU 시간 추정 방법을 활용.
  • '튜닝되지 않은' 모델을 명확하고 일관되게 정의할 것을 주장하여, 하이퍼파rameter 튜닝 효과의 간섭을 줄이고 공정한 벤치마크 비교를 가능하게 하기.

실험 결과

연구 질문

  • RQ12018–2019년도 주요 사전 훈련된 언어 모델들 간의 핵심 아키텍처적 및 자원 기반의 차이점은 무엇인가요?
  • RQ2트랜스포머, 양방향 어텐션 등의 아키텍처 혁신이 모델 규모나 사전 훈련 데이터 크기 대비 성능 향상에 얼마나 기여하는가요?
  • RQ3모델 크기, 계산 자원, 하이퍼파rameter 튜닝 등 중복 요소들이 성능 향상에 기여할 경우, 왜 모델 간의 공정한 비교가 어려운가요?
  • RQ4표준화된 보고 방식은 NLP 연구의 재현 가능성과 투명성을 어떻게 향상시킬 수 있나요?
  • RQ5광범위한 하이퍼파rameter 튜닝과 대규모 모델 훈련이 사전 훈련된 언어 모델에 미치는 환경적 및 실용적 영향은 무엇인가요?

주요 결과

  • NLP 논문들 간에 훈련 시간, 하이퍼파rameter 튜닝, 계산 자원 등의 모델 개발 세부 정보를 일관되게 보고하는 표준이 존재하지 않습니다.
  • 모델 아키텍처와 파라미터 수는 널리 보고되지만, 튜닝 방법, 튜닝 시간, 실험 실행 시간에 대한 세부 정보는 자주 생략되어 있어 재현 가능성을 떨어뜨립니다.
  • 성능과 재현 가능성에 결정적인 역할을 하는 어휘 자원(예: 사전 훈련 코퍼스)의 가용성 및 기술이 자주 누락되어 있습니다.
  • 특히 '튜닝되지 않은' 모델 성능과 계산 비용에 대해 글로벌로 인정받는 보고 프레임워크가 강력히 필요하며, 이는 공정한 벤치마크 비교를 가능하게 합니다.
  • 훈련 시간과 계산 자원 보고는 실현 가능하며 재현 가능성과 매우 관련성이 높지만, 현재는 거의 전면적으로 연구 논문에서 생략되어 있습니다.
  • 저자들은 실현 가능성과 영향력을 균형 잡은 계층적 보고 기준을 제안하며, '모델 아키텍처', '파라미터 수', '튜닝된 모델의 벤치마크 성능'은 핵심적이고 실현 가능한 보고 항목으로 간주합니다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.