Skip to main content
QUICK REVIEW

[논문 리뷰] CARE: a Benchmark Suite for the Classification and Retrieval of Enzymes

Jason Yang, Ariane Mora|arXiv (Cornell University)|2024. 06. 21.
Advanced Proteomics Techniques and Applications인용 수 7
한 줄 요약

CARE은 효소 기능에 대한 벤치마크 세트를 도입합니다. Task 1은 EC 번호로 효소 서열을 분류하고, Task 2는 반응으로부터 EC 번호를 검색합니다. 또한 다중모달 검색을 위한 기본 모델 CREEP이 포함됩니다.

ABSTRACT

Enzymes are important proteins that catalyze chemical reactions. In recent years, machine learning methods have emerged to predict enzyme function from sequence; however, there are no standardized benchmarks to evaluate these methods. We introduce CARE, a benchmark and dataset suite for the Classification And Retrieval of Enzymes (CARE). CARE centers on two tasks: (1) classification of a protein sequence by its enzyme commission (EC) number and (2) retrieval of an EC number given a chemical reaction. For each task, we design train-test splits to evaluate different kinds of out-of-distribution generalization that are relevant to real use cases. For the classification task, we provide baselines for state-of-the-art methods. Because the retrieval task has not been previously formalized, we propose a method called Contrastive Reaction-EnzymE Pretraining (CREEP) as one of the first baselines for this task and compare it to the recent method, CLIPZyme. CARE is available at https://github.com/jsunn-y/CARE/.

연구 동기 및 목표

  • EC 번호에 따른 효소 서열 분류와 반응으로 주어진 EC 번호를 검색하는 두 가지 현실적인 효소 기능 과제를 형식化합니다.
  • 단백질 서열과 EC를 연결하는 고품질 데이터셋과 반응과 EC를 연결하는 데이터셋을 선별하고, 분포 외 일반화를 위한 학습-테스트 분할을 제공합니다.
  • Task 1의 기준선을 제공하고 Task 2의 초기 기준선으로 CREEP를 도입합니다.
  • 다중 모달에 걸친 효소 기능 예측 및 검색에서 최신 모델의 벤치마킹을 가능하게 합니다.

제안 방법

  • 두 가지 작업(Task 1: EC에 따른 효소 서열 분류; Task 2: 반응으로부터 EC 검색)과 이에 해당하는 평가 설정을 정의합니다.
  • 단백질-EC 매핑을 위한 Swiss-Prot/UniProt 데이터셋과 반응-EC 매핑을 위한 EnzymeMap/ECReact 데이터셋을 선별합니다.
  • 두 작업 모두에 대해 도메인 외 일반화를 시뮬레이션하기 위한 학습-테스트 분할을 설계합니다(동일성 정도나 난이도 수준을 변동).
  • Task 1에서 기존 모델(CLEAN, BLASTp, ProteInfer 등)의 기준선 평가를 수행합니다.
  • Task 2 검색을 위해 반응(rxnfp)과 단백질(ProtT5)을 정렬하기 위한 다중모달 대조 학습(pretraining) 접근법인 CREEP를 제안하며, 선택적으로 텍스트 모달리티(SciBERT)를 포함할 수 있습니다.
  • CARE 저장소에서 오픈소스 벤치마크 리소스를 제공합니다.
Figure 1: Overview of CARE. (A) Dataset format for CARE, showing examples of enzymes and their associated reactions. The EC number acts as a bridge between a protein sequence and the reactions it is likely to perform. The EC number is a hierarchical classification scheme for enzyme function with fou
Figure 1: Overview of CARE. (A) Dataset format for CARE, showing examples of enzymes and their associated reactions. The EC number acts as a bridge between a protein sequence and the reactions it is likely to perform. The EC number is a hierarchical classification scheme for enzyme function with fou

실험 결과

연구 질문

  • RQ1모델이 학습 세트와 다른 도메인에 속하는 다양한 유사도의 서열을 다루는 경우 EC 번호로 효소 서열을 얼마나 잘 분류할 수 있을까요?
  • RQ2다중 모달 표현(단백질 서열, 반응 표현, 텍스트 설명)을 활용하여 본 적이 없는 반응에 대해 EC 번호를 검색할 수 있을까요?
  • RQ3새로운 효소-반응 쌍의 검색 성능은 다중 모달(텍스트 설명)을 도입함으로써 향상될까요?

주요 결과

  • 최첨단 분류기(CLEAN 등)가 Task 1에서 여러 분할에 걸쳐 무작위 베이스라인 및 조잡한 BLASTp 베이스라인을 상회합니다.
  • 일부 낮은 동정성 구간에서 BLASTp가 여전히 경쟁력이 있어 서열 유사성 기반의 벤치마크 가치가 강조됩니다.
  • 열려 있는 어휘(본 적 없는 EC)와 다중 모달성 때문에 Task 2는 더 어렵고, 쉬운 분할에서 어려운 분할로 갈수록 검색 성능이 감소합니다.
  • CREEP은 Task 2에서 강력한 기준선을 제공하며, 특히 텍스트 설명으로 보강될 때 다중 모달 대조 정렬의 이점을 얻습니다.
  • 더 어려운 분할에 걸쳐 대부분의 방법은 개선 여지가 크며, 다중 모달 및 고급 표현 전략의 필요성을 강조합니다.
Figure 2: Distribution of similarities between samples in each test set and the corresponding train set. (A) Protein sequence identity (Task 1) was measured to the closest hit in the train set using BLASTp. Sequence identity can be thought of as normalized Levenshtein distance. (B) Reaction similari
Figure 2: Distribution of similarities between samples in each test set and the corresponding train set. (A) Protein sequence identity (Task 1) was measured to the closest hit in the train set using BLASTp. Sequence identity can be thought of as normalized Levenshtein distance. (B) Reaction similari

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.