Skip to main content
QUICK REVIEW

[논문 리뷰] Protein Design with Agent Rosetta: A Case Study for Specialized Scientific Agents

Jacopo Teneggi, S. M. Bargeen A. Turzo|arXiv (Cornell University)|2026. 03. 16.
Machine Learning in Materials Science인용 수 0
한 줄 요약

Agent Rosetta를 소개합니다. RosettaScripts와 통합된 LLM 주도 에이전트로 단백질을 반복적으로 설계하며 ML 기준선과 경쟁력 있는 성능을 달성하고 비정형 잔류물(non-canonical residues)을 가능하게 합니다.

ABSTRACT

Large language models (LLMs) are capable of emulating reasoning and using tools, creating opportunities for autonomous agents that execute complex scientific tasks. Protein design provides a natural testbed: although machine learning (ML) methods achieve strong results, these are largely restricted to canonical amino acids and narrow objectives, leaving unfilled need for a generalist tool for broad design pipelines. We introduce Agent Rosetta, an LLM agent paired with a structured environment for operating Rosetta, the leading physics-based heteropolymer design software, capable of modeling non-canonical building blocks and geometries. Agent Rosetta iteratively refines designs to achieve user-defined objectives, combining LLM reasoning with Rosetta's generality. We evaluate Agent Rosetta on design with canonical amino acids, matching specialized models and expert baselines, and with non-canonical residues -- where ML approaches fail -- achieving comparable performance. Critically, prompt engineering alone often fails to generate Rosetta actions, demonstrating that environment design is essential for integrating LLM agents with specialized software. Our results show that properly designed environments enable LLM agents to make scientific software accessible while matching specialized tools and human experts.

연구 동기 및 목표

  • 자율 에이전트를 사용하여 Rosetta 기반의 복잡한 단백질 설계 업무를 자동화할 것을 동기 부여한다.
  • 환경 설계가 LLM과 도메인 특화 소프트웨어를 연결하는 데 필수적임을 보여준다.
  • Agent Rosetta가 정형 설계에서 특화된 ML 모델에 필적하고 비정형 잔류물에서 인간 기준선을 능가할 수 있음을 보여준다.
  • 중간 설계 메트릭에 따라 프로토콜을 조정하는 다중 턴 최적화 프레임워크를 제공한다.

제안 방법

  • Robust한 액션 생성을 위한 Tailored RosettaScripts 환경과 인터페이스가 있는 LLM 에이전트인 Agent Rosetta를 개발한다.
  • Surrogate 메트릭(회전 반경(radius of gyration), 공동공간 부피(cavity volume), 매립된 불충분 수소결합(buried unsatisfied H-bonds), 목표에 대한 RMSD, pLDDT)을 사용하여 Pose 데이터 전체를 context 내에서 필요로 하지 않으면서 의사 결정을 유도한다.
  • 세 가지 액션 유형(rotamer_change, backbone_change, go_back_to)을 정의하여 RosettaScripts 내의 설계 파이프라인을 포괄한다.
  • 구조화된 다중 턴 추론을 사용: 액션 선택, 환경 문서를 통한 매개변수 생성, 실행 및 피드백을 통한 설계 개선.
  • 의미적으로 올바른 액션을 보장하고 안정적인 다중 턴 상호작용을 가능하게 하기 위해 RosettaScripts의 페널티와 연산을 간소화된 템플릿으로 추상화한다.
Figure 1 : Illustration of our multi-turn agentic system. (A) Schematics of Agent Rosetta’s interaction protocol. (B) Design refinement: the agent chooses the action, and, after the environment returns the action documentation, it generates the action call with its parameters.
Figure 1 : Illustration of our multi-turn agentic system. (A) Schematics of Agent Rosetta’s interaction protocol. (B) Design refinement: the agent chooses the action, and, after the environment returns the action documentation, it generates the action call with its parameters.

실험 결과

연구 질문

  • RQ1LLM 에이전트가 구조화된 환경을 통해 Rosetta를 효과적으로 제어하여 단백질 설계 목표를 최적화할 수 있는가?
  • RQ2환경 설계(프롬프트를 넘어서)가 RosettaScripts를 위한 도메인 특화 액션의 안정적인 생성을 가능하게 하는가?
  • RQ3Canonically 고정 백본 설계에서 Agent Rosetta는 ProteinMPNN 및 인간 프로토콜과 비교해 어떤 성과를 보이는가?
  • RQ4Agent Rosetta가 비정형 아미노산을 단백질 코어에 설계할 수 있는가? 이는 ML 모델들이 어려운 영역이다?

주요 결과

  • Agent Rosetta는 Canonical 아미노산에서 ProteinMPNN과 비교할 만한 설계 품질을 달성한다(0.20 Å RMSD 허용 오차 이내).
  • Agent Rosetta는 비정형 아미노산 작업에서 전문가 인간 기준선을 능가하여 데이터 기반 방법을 넘어서는 능력을 보인다.
  • 프롬프트만으로는 Rosetta 액션을 안정적으로 생성하기에 충분하지 않으며, 강건한 에이전트 성능을 위해 환경 및 구문 설계가 필수적이다.
  • 맞춤형 환경으로 모든 테스트된 LLM은 액션 성공률이 ≥ 86%를 달성했다.
  • GPT-5 기반 구성이 여러 실험에서 최상의 비용-성능 트레이드오프를 보여주었다.
  • NCAA 설계에서 Agent Rosetta는 인간 기준선보다 높은 AF3 RMSD 및 pLDDT를 달성하여 구조 검증이 평균적으로 개선되었음을 시사한다.
Figure 2 : A failure example of prompting for generation of composition penalties. Even though Agent Rosetta wants to reduce proline content, the penalty block achieves the opposite effect.
Figure 2 : A failure example of prompting for generation of composition penalties. Even though Agent Rosetta wants to reduce proline content, the penalty block achieves the opposite effect.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.