Skip to main content
QUICK REVIEW

[논문 리뷰] DeepSpeed4Science Initiative: Enabling Large-Scale Scientific Discovery through Sophisticated AI System Technologies

Shuaiwen Leon Song, Bonnie Kruft|arXiv (Cornell University)|2023. 10. 06.
Machine Learning in Materials Science인용 수 8
한 줄 요약

이 논문은 DeepSpeed4Science 이니셔티브를 소개하며, 대규모 과학 발견을 가속화하기 위해 DeepSpeed에 구축된 AI 시스템 기술들을 상세히 제시하고, 두 개의 구조생물학 시연과 더 넓은 과학 협력을 위한 계획을 제시한다.

ABSTRACT

In the upcoming decade, deep learning may revolutionize the natural sciences, enhancing our capacity to model and predict natural occurrences. This could herald a new era of scientific exploration, bringing significant advancements across sectors from drug development to renewable energy. To answer this call, we present DeepSpeed4Science initiative (deepspeed4science.ai) which aims to build unique capabilities through AI system technology innovations to help domain experts to unlock today's biggest science mysteries. By leveraging DeepSpeed's current technology pillars (training, inference and compression) as base technology enablers, DeepSpeed4Science will create a new set of AI system technologies tailored for accelerating scientific discoveries by addressing their unique complexity beyond the common technical approaches used for accelerating generic large language models (LLMs). In this paper, we showcase the early progress we made with DeepSpeed4Science in addressing two of the critical system challenges in structural biology research.

연구 동기 및 목표

  • 일반 LLM 가속화 너머의 과학 도메인에 맞춘 AI 시스템 기술의 필요성에 대한 동기 부여.
  • DeepSpeed4Science 접근 방식과 그 기반이 되는 DeepSpeed 기둥(훈련, 추론, 압축)의 설명.
  • DS4Sci가 다루는 구조생물학의 두 가지 시스템 도전 과제 제시( Evoformer 주의에서의 메모리 폭주; GenSLMs를 위한 긴 시퀀스 지원).
  • 협업 모델 및 과학 지향 AI 시스템 기술 공유를 위한 잠재 플랫폼의 개요.

제안 방법

  • 메모리 폭주를 제거하기 위해 메모리 효율적인 EvoformerAttention 커널을 개발.
  • 메가트론-딥스피드 프레임워크를 긴 시퀀스 지원으로 강화하고 주의 마스크 및 위치 임베딩에 대한 메모리 최적화를 통합.
  • 메가트론-딥스피드 재베이스를 통해 게놈 규모 기초 모델의 매우 긴 시퀀스 훈련/추론 가능.
  • 커널 융합 및 타일링, 즉시 브로드캐스팅, FP32-안전 그래디언트 처리를 통해 피크 메모리를 줄이면서 정확도를 유지.
  • 시퀀스 병렬성, 텐서/파이프라인 병렬성, 모델/데이터 오프로드를 활용해 시퀀스 길이를 극적으로 확장.
Figure 1: DeepSpeed4Science approach: developing a new set of AI system technologies that are beyond generic large language model support, tailored for accelerating scientific discoveries and addressing their complexity.
Figure 1: DeepSpeed4Science approach: developing a new set of AI system technologies that are beyond generic large language model support, tailored for accelerating scientific discoveries and addressing their complexity.

실험 결과

연구 질문

  • RQ1과학 중심 모델의 메모리 및 시퀀스 길이 문제를 해결하기 위해 AI 시스템 기술을 어떻게 특화할 수 있는가?
  • RQ2맞춤형 커널과 프레임워크 재배치를 통해 게놈 규모 및 Evoformer 기반 모델의 컨텍스트 크기를 정확도 저하 없이 훨씬 더 길게 만들 수 있는가?
  • RQ3DS4Sci 최적화를 구조생물학 및 GenSLM 스타일 모델에 적용했을 때 성능/처기량 향상은 어느 정도인가?
  • RQ4DS4Sci가 과학적 발견을 위한 첨단 AI 시스템 기술의 더 넓은 협업과 공유를 어떻게 촉진할 수 있는가?

주요 결과

  • DS4Sci_EvoformerAttention 커널이 OpenFold의 피크 메모리를 Evoformer-attention 변형에서 정확도 손실 없이 13배 감소시켰다.
  • 새로운 Megatron-DeepSpeed 프레임워크가 GenSLMs를 더 긴 시퀀스로 학습 가능하게 하였으며, 평균적으로 최대 13배 더 긴 시퀀스와 특정 경우 최대 2배 처리량을 보고한다.
  • Megatron-DeepSpeed 재베이스에 로터리 위치 임베딩, FlashAttention v1/v2 및 새로운 융합 커널 추가로 긴 시퀀스 학습과 추론 개선.
  • 주목 마스크 및 위치 임베딩의 메모리 최적화와 시퀀스 병렬성으로 GenSLMs의 실행 가능한 시퀀스 길이가 크게 확장(예: 25B GenSLM의 512K) 이전 한계를 넘겼다.
  • DS4Sci 노력은 DeepSpeed4Science를 과학를 위한 고급 AI 시스템 기술 공유를 위한 플랫폼이자 저장소로 위치시킨다.
Figure 2: Peak memory requirement for training variants of the MSA attention kernels (with bias) with the maximum possible training sample dimension in OpenFold. (Left) The original OpenFold implementation with EvoformerAttention used in AlphaFold2. The memory explosion problems in training/inferenc
Figure 2: Peak memory requirement for training variants of the MSA attention kernels (with bias) with the maximum possible training sample dimension in OpenFold. (Left) The original OpenFold implementation with EvoformerAttention used in AlphaFold2. The memory explosion problems in training/inferenc

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.