Skip to main content
QUICK REVIEW

[논문 리뷰] Masked Language Modeling for Proteins via Linearly Scalable Long-Context Transformers

Krzysztof Choromański, Valerii Likhosherstov|arXiv (Cornell University)|2020. 06. 05.
Machine Learning in Bioinformatics인용 수 40
한 줄 요약

Performer를 도입한 Transformer 변형으로, Fast Attention Via Orthogonal Random features (FAVOR)를 사용하여 선형 스케일의 어텐션을 달성하고, 이로써 이론적 보장 및 역호환성과 함께 긴 컨텍스트 단백질 서열 모델링이 가능해진다.

ABSTRACT

Transformer models have achieved state-of-the-art results across a diverse range of domains. However, concern over the cost of training the attention mechanism to learn complex dependencies between distant inputs continues to grow. In response, solutions that exploit the structure and sparsity of the learned attention matrix have blossomed. However, real-world applications that involve long sequences, such as biological sequence analysis, may fall short of meeting these assumptions, precluding exploration of these models. To address this challenge, we present a new Transformer architecture, Performer, based on Fast Attention Via Orthogonal Random features (FAVOR). Our mechanism scales linearly rather than quadratically in the number of tokens in the sequence, is characterized by sub-quadratic space complexity and does not incorporate any sparsity pattern priors. Furthermore, it provides strong theoretical guarantees: unbiased estimation of the attention matrix and uniform convergence. It is also backwards-compatible with pre-trained regular Transformers. We demonstrate its effectiveness on the challenging task of protein sequence modeling and provide detailed theoretical analysis.

연구 동기 및 목표

  • 희소성 선행 가정에 의존하지 않고, 긴 생물학적 서열(예: 단백질)에 맞춰 Transformer 모델의 규모 확장을 촉진한다.
  • 이론적 보장과 학 pretrained Transformer와의 실용적 호환성을 갖춘 선형 시간 어텐션 메커니즘을 개발한다.
  • 단백질 서열 모델링과 ImageNet64에서의 효과를 입증하면서 복잡도와 수렴성을 분석한다.

제안 방법

  • 커널 기반일 수 있는 Generalized Attention (GA)을 제시한다.
  • 일반 어텐션을 Fast Attention Via Orthogonal Random features (FAVOR)로 교체하여 어텐션 행렬의 편향되지 않은 저랭크 근사를 얻는다.
  • 쿼리(query)와 키(key) 사이의 커널 유사도를 추정하기 위해 무작위 특징 맵을 사용하고, 전체 A 행렬을 형성하지 않고도 Q'K'ᵀ의 계산 형태를 얻는다.
  • 분산을 줄이고 근사 품질을 향상시키기 위해 Orthogonal Random Features (ORFs)를 도입한다.
  • O(Ld log d) 공간과 O(Ld² log d) 시간의 시간/공간 복잡도 분석을 제시하고, 일반 어텐션의 O(L²d)와 비교한다.
  • 다른 Transformer 구성 요소를 손대지 않고 어텐션 메커니즘만 교체하여 역호환성을 보여준다.

실험 결과

연구 질문

  • RQ1FAVOR가 일반화된 커널 기반 어텐션에 대해 편향되지 않은 추정으로 표준 어텐션 행렬을 근사할 수 있는가?
  • RQ2FAVOR 기반 어텐션이 길이가 L인 서열에서 선형적으로 확장되면서 긴 단백질 서열에 대한 정확도를 유지하는가?
  • RQ3Performer는 일반 Transformer 및 희소 어텐션 방법과 비교하여 긴 형식의 단백질 서열 모델링과 ImageNet64에서 어떤 성능을 보이는가?
  • RQ4GA 프레임워크에서 FAVOR의 이론적 보장(편향성 및 균일 수렴)에 대한 보장은 무엇인가?

주요 결과

  • FAVOR는 토큰 수에 대해 선형 시간과 저차원 공간을 달성하며, 제곱 어텐션을 저랭크 근사로 대체한다.
  • 이 방법은 어텐션 행렬의 편향 없는 추정과 균일 수렴 보장을 제공한다.
  • FAVOR로 학습된 Performers는 미세 조정을 거쳐 단백질 서열 작업에서 Transformer 성능을 회복하거나 근접하게 매칭하며, 단백질에 대한 일부 희소 어텐션 기반 방법들보다 우수할 수 있다.
  • Orthogonal random features는 비구조적 특징에 비해 근사 오차를 감소시키고 하류 성능을 향상시킨다.
  • 이 방법은 사전 학습된 일반 Transformer와 API 호환 가능하며, 일반 어텐션의 확 scalable 대체로 사용될 수 있다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.