[論文レビュー] Protein-DNA binding sites prediction based on pre-trained protein language model and contrastive learning
本稿では、事前学習されたタンパク質言語モデルと対照的学習を統合することで、タンパク質-DNA結合部位を予測する新しい深層学習フレームワーク、CLAPE-DBを提案する。自己教師あり表現学習とタンパク質配列における対照的事前学習を活用することで、CLAPE-DBは2つのベンチマークデータセットでAUCスコア0.871および0.881を達成し、既存手法よりも精度と一般化性能に優れている。
Protein-DNA interaction is critical for life activities such as replication, transcription, and splicing. Identifying protein-DNA binding residues is essential for modeling their interaction and downstream studies. However, developing accurate and efficient computational methods for this task remains challenging. Improvements in this area have the potential to drive novel applications in biotechnology and drug design. In this study, we propose a novel approach called CLAPE, which combines a pre-trained protein language model and the contrastive learning method to predict DNA binding residues. We trained the CLAPE-DB model on the protein-DNA binding sites dataset and evaluated the model performance and generalization ability through various experiments. The results showed that the AUC values of the CLAPE-DB model on the two benchmark datasets reached 0.871 and 0.881, respectively, indicating superior performance compared to other existing models. CLAPE-DB showed better generalization ability and was specific to DNA-binding sites. In addition, we trained CLAPE on different protein-ligand binding sites datasets, demonstrating that CLAPE is a general framework for binding sites prediction. To facilitate the scientific community, the benchmark datasets and codes are freely available at https://github.com/YAndrewL/clape.
研究の動機と目的
- タンパク質-DNA結合部位をより正確かつ一般化可能な計算手法で予測するための開発。
- タンパク質-DNA相互作用におけるラベル付きデータの不足と複雑な配列-構造関係の課題に対処する。
- 自己教師あり事前学習と対照的学習を活用して、結合部位予測のための表現学習を強化する。
- DNA結合にとどまらず、他のリガンド結合部位にも適用可能な汎用フレームワークの構築。
提案手法
- 本手法は、タンパク質配列を文脈に応じた埋め込み表現に変換するため、事前学習済みタンパク質言語モデル(例:ESMまたは類似モデル)を用いる。
- 対照的学習を適用し、正例と負例のシーケンスビューを対比させることで、判別可能な表現を学習する。
- 正例ペアはデータ拡張(例:ランダムマスキングやノイズ注入)により生成され、負例ペアは異なるタンパク質配列から形成される。
- モデルは、タンパク質-DNA結合データセット上でエンドツーエンドに微調整され、結合部位の予測が行われる。
- 対照的損失関数により、トレーニング中に正例ペアを引き寄せ、負例ペアを遠ざけるよう促進される。
- 最終的なモデル、CLAPE-DBは、AUCやAUPRCなどの標準指標を用いてベンチマークデータセットで評価される。
実験結果
リサーチクエスチョン
- RQ1対照的学習は、DNA結合部位の予測におけるタンパク質配列の表現学習を改善できるか?
- RQ2標準ベンチマークデータセットにおいて、CLAPE-DBは既存の最先端モデルと比較してどの程度の性能を示すか?
- RQ3本フレームワークは、DNA結合にとどまらず、他の種類のリガンド結合部位に対しても一般化できるか?
- RQ4この文脈において、事前学習されたタンパク質言語モデルと対照的学習を組み合わせた際の寄与度は何か?
- RQ5CLAPE-DBは、配列データの変動やデータスパarsityに対してどの程度のロバストネスを示すか?
主な発見
- CLAPE-DBは最初のベンチマークデータセットでAUC 0.871を達成し、優れた予測性能を示した。
- 2番目のベンチマークデータセットでは、CLAPE-DBはAUC 0.881に達し、高い精度とロバストネスを示した。
- 限られたラベル付きデータでも、既存手法と比較して優れた一般化能力を示した。
- フレームワークは、タンパク質-リガンド結合部位の予測に対しても成功裏に適応され、異なる結合タイプへの一般化可能性が確認された。
- 事前学習された言語モデルと対照的学習の統合により、結合部位予測のための表現学習が顕著に向上した。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。