[논문 리뷰] A Multi-task Large Reasoning Model for Molecular Science
논문은 다-task 대규모 추론 모델을 다중 전문화 아키텍처와 사슬 사고(chain-of-thought) 추론으로 강화하고 강화학습으로 개선하여 데이터 효율적 학습으로 강력한 다-task 분자 성능을 달성한다.
Advancements in artificial intelligence for molecular science are necessitating a paradigm shift from purely data-driven predictions to knowledge-guided computational reasoning. Existing molecular models are predominantly proprietary, lacking general molecular intelligence and generalizability. This underscores the necessity for computational methods that can effectively integrate scientific logic with deep learning architectures. Here we introduce a multi-task large reasoning model designed to emulate the cognitive processes of molecular scientists through structured reasoning and reflection. Our approach incorporates multi-specialist modules to provide versatile molecular expertise and a chain-of-thought (CoT) framework enhanced by reinforcement learning infused with molecular knowledge, enabling structured and reflective reasoning. Systematic evaluations across 10 molecular tasks and 47 metrics demonstrate that our model achieves an average 50.3% improvement over the base architecture, outperforming over 20 state-of-the-art baselines, including ultra-large-parameter foundation models, despite using significantly fewer training data and computational resources. This validates that embedding explicit reasoning mechanisms enables high-efficiency learning, allowing smaller-scale models to surpass massive counterparts in both efficacy and interpretability. The practical utility of this computational framework was validated through a case study on the design of central nervous system (CNS) drug candidates, illustrating its capacity to bridge data-driven and knowledge-integrated approaches for intelligent molecular design.
연구 동기 및 목표
- 순수 예측을 넘어 분자 작업에 화학 지식을 딥러닝에 통합하도록 동기를 부여한다.
- 화학 로직을 CoT 추론에 내재화하는 다중 전문화, 작업 적응형 프레임워크를 개발한다.
- 데이터 시너지와 전문가 시너지를 강화학습과 결합하여 데이터 효율적인 학습을 달성한다.
- 제한된 학습 데이터와 자원으로 열 가지 분자 작업에서 우수한 다중 과제 성능을 입증한다.
- 생성, 예측 및 합성을 연결하는 CNS 약물 설계 사례 연구를 통해 실용적 활용성을 보여준다.
제안 방법
- 사전 학습된 LLM(DeepSeek-7B 기반) 내부에 다중 전문화 계층을 구성하고 작업 유형별로 여덟 개 전문 그룹을 조정하는 라우터를 도입한다.
- 예측 전문가를 93K 지시 데이터셋에서 학습시키고 추론(CoT) 전문가는 3.5K 개의 고품질 CoT 데이터셋에서 학습시킨다.
- Low-Rank Adaptation(LoRA)을 도입하여 매개변수 업데이트의 효율성을 높인다.
- 작업 특화 분자 과학 보상으로 강화학습을 적용하여 추론을 화학적 타당성과 일치시킨다.
- 3단계 학습을 사용: 74.5K 데이터에서 지시 미세조정을 통한 표현 학습, 3.6K 데이터에서 CoT 미세조정, 지식 정렬 RL.
- 관련 작업의 공동 학습인 데이터 시너지와 예측+추론 전문가 협력인 전문시너지로 추론을 향상시킨다.

실험 결과
연구 질문
- RQ1다중-task 분자 추론 모델이 화학 지식을 체인 오브 생각 추론에 내재화하여 다양한 작업에서 최첨단 베이스라인을 능가할 수 있는가?
- RQ2데이터 시너지와 전문가 시너지가 다중-task 분자 성능 및 추론 정렬에 어떤 영향을 미치는가?
- RQ3지식 유도 보상으로의 강화학습이 예측 전문가와 추론 전문가 간의 일관성에 어떤 영향을 주는가?
- RQ4CNS 약물 설계 시나리오에서 지식이 주입된 더 작고 해석 가능한 모델로 높은 정확도와 해석 가능한 추론을 제공하는 것이 가능한가?
주요 결과
- 10개의 분자 작업에서 기준 아키텍처 대비 평균 50.3%의 개선.
- 훈련 데이터 및 자원을 덜 사용하면서 20개 이상의 최첨단 베이스라인을 능가하며, 초대형 매개변수 모델 포함.
- 강력한 다중 작업 모델 LLaSMol에 비해 작업 지표에서 약 6%의 개선에 근접.
- 사슬 사고를 통한 견고한 추론 해석 가능성과 CNS 약물 설계 사례 연구를 통해 입증.
- 데이터 시너지와 전문가 시너지 및 CoT RL이 지시어만 또는 CoT만Variant 대비 성능을 크게 향상시킴.
- 모델이 기저선보다 약간 뒤지는 태스크로 Lipophilicity를 확인, 전문화 한계 시사

더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.