[논문 리뷰] Mol-Instructions: A Large-Scale Biomolecular Instruction Dataset for Large Language Models
Mol-Instructions는 LLM들을 위한 대규모 생체분자 지시 데이터셋으로, 분자-, 단백질-, 생체분자 텍스트 지향 작업을 다루며, 여러 모델에서 지시 조정을 통한 개선을 보여줍니다. 지속적인 연구를 위해 공개적으로 이용 가능하며 정기적으로 업데이트됩니다.
Large Language Models (LLMs), with their remarkable task-handling capabilities and innovative outputs, have catalyzed significant advancements across a spectrum of fields. However, their proficiency within specialized domains such as biomolecular studies remains limited. To address this challenge, we introduce Mol-Instructions, a comprehensive instruction dataset designed for the biomolecular domain. Mol-Instructions encompasses three key components: molecule-oriented instructions, protein-oriented instructions, and biomolecular text instructions. Each component aims to improve the understanding and prediction capabilities of LLMs concerning biomolecular features and behaviors. Through extensive instruction tuning experiments on LLMs, we demonstrate the effectiveness of Mol-Instructions in enhancing large models' performance in the intricate realm of biomolecular studies, thus fostering progress in the biomolecular research community. Mol-Instructions is publicly available for ongoing research and will undergo regular updates to enhance its applicability.
연구 동기 및 목표
- LLMs를 위한 격차를 메우기 위해 전용 biomolecular 지시 데이터셋의 생성을 촉진한다.
- 세 가지 핵심 구성 요소인 molecule-oriented, protein-oriented, 및 biomolecular text 지시로 Mol-Instructions를 구성한다.
- Mol-Instructions를 활용한 LLM들에 대한 지시 조정의 효과를 여러 baselines에서 시연한다.
- 데이터셋에 대한 공개 접근성을 제공하고 적용 가능 범위를 넓히기 위한 향후 개선 사항을 제시한다.
제안 방법
- 자가 설명(self-instruct), 템플릿 기반 변환, 인간이 제작한 설명의 혼합 구성 방법을 통해 세 도메인에 걸친 2백만 개가 넘는 생체분자 지시를 수집한다.
- 수동 품질 검사를 포함하여 GPT-3.5-turbo를 사용한 다양한 작업 설명 생성을 위해 인간–AI 협업을 구현한다.
- 표준 생화학 데이터베이스와 PubMed에서 데이터를 소싱하고, 데이터 마이닝 및 AI 지원 생성을 통해 입력/출력, QA 쌍, 설계 지침을 도출한다.
- 템플릿을 사용하여 UniProtKB 기반 단백질 설계 주석 등을 포함한 생물학 데이터를 텍스트 형식으로 변환하여 사용자가 지정한 목표를 충족시킨다.
- 엄격한 품질 관리 적용: 분자에 대해 SMILES를 SELFIES로 교체하고, UniProtKB 항목을 선별하며, 중복 감소를 위해 90% 유사도로 MMseqs로 단백질을 클러스터링한다.
- 훈련/검증/테스트 분할을 사용하여 세 가지 지시 도메인 전반에 걸쳐 LLama-7B 및 다른 baselines에서 지시 조정을 통한 평가를 실시한다.
실험 결과
연구 질문
- RQ1Mol-Instructions가 baselines와 비교하여 생체분자 이해 및 생성 작업에서 LLM 성능을 향상시키는가?
- RQ2분자-, 단백질-, 텍스트 지향 지시가 각 작업에서 어떠한 개선에 기여하는가?
- RQ3생성된 단백질 설계 및 분자 설명이 알려진 기능적 또는 구조적 주석과 일치하는가?
- RQ4데이터셋 구성 선택(자체 지시, 템플릿, 인간이 제작한 설명)이 모델 성능에 미치는 영향은 무엇인가?
주요 결과
- Mol-Instructions는 평가된 모델과 메트릭에서 baselines에 비해 분자 이해 작업에서 현저한 개선을 제공합니다.
- 데이터는 분자 특성 예측 및 생성 작업에서 향상된 성능을 가능하게 하며, 생성된 분자는 참조 구조와의 유사도가 더 높게 나타납니다.
- 단백질 관련 작업에서 조정된 모델은 기본적인 단백질 특징을 식별하고 de novo 설계를 UniProtKB 주석과 일치시킬 수 있음을 보여줘 기능적 관련성을 시사합니다.
- Mol-Instructions는 정보 추출 및 생물정보학 맥락의 Q&A를 포함한 생체분자 NLP 작업의 성능을 향상시킵니다.
- 도메인 특화 소형 모델과 비교할 때 Mol-Instructions로 학습된 대형 모델은 여전히 전문적인 생성에 격차를 보이지만 더 넓은 도메인 이해의 이점을 보입니다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.