Skip to main content
QUICK REVIEW

[논문 리뷰] MAP Format for Representing Chemical Modifications, Annotations, and Mutations in Protein Sequences: An Extension of the FASTA Format

Aditi Shendre, Naman Kumar Mehta|ArXiv.org|2025. 05. 06.
Genetics, Bioinformatics, and Biomedical Research인용 수 5
한 줄 요약

MAP은 헤더 메타 태그와 인라인 잔류 태그를 통해 화학 수정, 주석 및 돌연변이를 인코딩하기 위해 FASTA를 확장한 새로운 단백질 서열 형식입니다.

ABSTRACT

Several formats, including FASTA, PIR, GenBank, EMBL, and GCG, have been developed for representing protein sequences composed of natural amino acids. Among these, FASTA remains the most widely used due to its simplicity and human readability. However, FASTA lacks the capability to represent chemically modified or non-natural residues, as well as structural annotations and mutations in protein variants. To address some of these limitations, the PEFF format was recently introduced as an extension of FASTA. Additionally, formats such as HELM and BILN have been proposed to represent amino acids and their modifications at the atomic level. Despite their advancements, these formats have not achieved widespread adoption within the bioinformatics community due to their complexity. To complement existing formats and overcome current challenges, we propose a new format called MAP (Modification and Annotation in Proteins), which enables comprehensive annotation of protein sequences. MAP introduces meta tags in the header for protein-level annotations and inline tags within the sequence for residue-level modifications. In this format, standard one-letter amino acid codes are augmented with curly-brace tags to denote various modifications, including phosphorylation, acetylation, non-natural residues, cyclization, and other residue-specific features. The header metadata also captures information such as organism, function, and sequence variants. We describe the structure, objectives, and capabilities of the MAP format and demonstrate its application in bioinformatics, particularly in the domain of protein therapeutics. To facilitate community adoption, we are developing a comprehensive suite of MAP-format resources, including a detailed manual, annotated datasets, and conversion tools, available at http://webs.iiitd.edu.in/raghava/maprepo/.

연구 동기 및 목표

  • 자연 아미노산을 넘는 단백질 서열의 포괄적 주석 가능성을 실현한다.
  • 수정, 비자연 잔류물 및 돌연변이를 지원하는 실용적이고 읽기 쉬운 형식을 제공한다.
  • 더 간단하고 채택 친화적인 접근 방식으로 기존 형식(예: PEFF)을 보완한다.

제안 방법

  • 단백질 주석(생물체, 기능, 변종)에 대한 헤더 수준 메타 태그를 도입한다.
  • 잔류 수준의 수정 표시를 위해 표준 한 글자 아미노산 코드에 뒤에 붙는 중괄호 태그를 도입한다.
  • 지원하는 수정 유형(예: 인산화, 아세틸화, 비자연 잔류물, 고리화)을 설명한다.
  • MAP의 구조와 기능 및 의도된 생물정보학 응용을 설명한다.
  • 매뉴얼, 주석 데이터 세트 및 변환 도구와 같은 커뮤니티 리소스 계획을 개요한다.

실험 결과

연구 질문

  • RQ1MAP는 단백질 서열에서 잔류 수준의 수정들을 어떻게 인코딩하는가?
  • RQ2단백질 주석 및 변종에 필요한 또는 지원되는 헤더 메타데이터는 무엇인가?
  • RQ3MAP는 사용 편의성과 채택 가능성 측면에서 PEFF, HELM, BILN 같은 기존 형식과 어떻게 비교되는가?
  • RQ4MAP를 단백질 therapeutics 및 변이 주석과 같은 영역에 효과적으로 적용할 수 있는가?

주요 결과

  • MAP는 단백질 서열에서 수정 및 주석을 나타내기 위해 헤더 메타 태그와 인라인 잔류 수준의 중괄호 태그를 도입한다.
  • MAP는 인산화, 아세틸화, 비자연 잔류물 및 고리화 등을 포함한 잔류 특정 기능의 범위를 지원한다.
  • 이 형식은 구조, 목적 및 기능의 관점에서 설명되며 단백질 치료제와의 관련성이 시사된다.
  • 저자들은 더 넓은 MAP 생태계를 개발 중이며 상세 매뉴얼, 주석 데이터 세트 및 변환 도구를 포함한다.
  • MAP는 기존 형식의 한계를 해결하기 위한 FASTA의 보완적 확장으로 위치한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.