[论文解读] MAP Format for Representing Chemical Modifications, Annotations, and Mutations in Protein Sequences: An Extension of the FASTA Format
MAP 是一种新的蛋白质序列格式,扩展 FASTA,通过头部元标签和内联残基标签来编码化学修饰、注释和突变。
Several formats, including FASTA, PIR, GenBank, EMBL, and GCG, have been developed for representing protein sequences composed of natural amino acids. Among these, FASTA remains the most widely used due to its simplicity and human readability. However, FASTA lacks the capability to represent chemically modified or non-natural residues, as well as structural annotations and mutations in protein variants. To address some of these limitations, the PEFF format was recently introduced as an extension of FASTA. Additionally, formats such as HELM and BILN have been proposed to represent amino acids and their modifications at the atomic level. Despite their advancements, these formats have not achieved widespread adoption within the bioinformatics community due to their complexity. To complement existing formats and overcome current challenges, we propose a new format called MAP (Modification and Annotation in Proteins), which enables comprehensive annotation of protein sequences. MAP introduces meta tags in the header for protein-level annotations and inline tags within the sequence for residue-level modifications. In this format, standard one-letter amino acid codes are augmented with curly-brace tags to denote various modifications, including phosphorylation, acetylation, non-natural residues, cyclization, and other residue-specific features. The header metadata also captures information such as organism, function, and sequence variants. We describe the structure, objectives, and capabilities of the MAP format and demonstrate its application in bioinformatics, particularly in the domain of protein therapeutics. To facilitate community adoption, we are developing a comprehensive suite of MAP-format resources, including a detailed manual, annotated datasets, and conversion tools, available at http://webs.iiitd.edu.in/raghava/maprepo/.
研究动机与目标
- 实现对蛋白质序列超出天然氨基酸的全面注释。
- 提供一个实用、易读的格式,支持修饰、非天然残基和突变。
- 用更简单、便于采用的方法补充现有格式(如 PEFF)。
提出的方法
- 引入用于蛋白质注释的头部级元标签(生物来源、功能、变体)。
- 在标准单字母氨基酸代码后附加花括号内联标签,以表示残基级别修饰。
- 描述支持的修饰类型(例如,磷酸化、乙酰化、非天然残基、环化)。
- 解释 MAP 的结构与能力及其面向的生物信息学应用。
- 概述社区资源计划:手册、带注释的数据集和转换工具。
实验结果
研究问题
- RQ1MAP 如何在蛋白质序列中对残基级修饰进行编码?
- RQ2用于蛋白质注释和变体的头部元数据有哪些是必需的或受支持的?
- RQ3在可用性和采用潜力方面,MAP 与现有格式(如 PEFF、HELM、BILN)相比有何差异?
- RQ4MAP 是否能有效应用于蛋白质治疗和变体注释等领域?
主要发现
- MAP 引入头部元标签和内联残基级花括号标签,以表示蛋白质序列中的修饰和注释。
- MAP 支持一系列残基特定特征,包括磷酸化、乙酰化、非天然残基以及环化。
- 该格式在结构、目标和能力方面进行描述,并在蛋白质治疗方面具有相关性证明。
- 作者正在开发更广泛的 MAP 生态系统,包括详细手册、带注释的数据集和转换工具。
- MAP 被定位为对 FASTA 的互补扩展,以解决现有格式的局限性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。