[논문 리뷰] Fengshenbang 1.0: Being the Foundation of Chinese Cognitive Intelligence
Open-source Chinese foundation-model ecosystem (Fengshenbang)을 세 가지 구성 요소—Models, Framework, Benchmark—와 함께 제안하고, 49-model 카탈로그와 중국 특화 평가 벤치마크를 통해 접근 가능하고 자원 효율적인 중국어 NLP 개발을 가능하게 한다.
Nowadays, foundation models become one of fundamental infrastructures in artificial intelligence, paving ways to the general intelligence. However, the reality presents two urgent challenges: existing foundation models are dominated by the English-language community; users are often given limited resources and thus cannot always use foundation models. To support the development of the Chinese-language community, we introduce an open-source project, called Fengshenbang, which leads by the research center for Cognitive Computing and Natural Language (CCNL). Our project has comprehensive capabilities, including large pre-trained models, user-friendly APIs, benchmarks, datasets, and others. We wrap all these in three sub-projects: the Fengshenbang Model, the Fengshen Framework, and the Fengshen Benchmark. An open-source roadmap, Fengshenbang, aims to re-evaluate the open-source community of Chinese pre-trained large-scale models, prompting the development of the entire Chinese large-scale model community. We also want to build a user-centered open-source ecosystem to allow individuals to access the desired models to match their computing resources. Furthermore, we invite companies, colleges, and research institutions to collaborate with us to build the large-scale open-source model-based ecosystem. We hope that this project will be the foundation of Chinese cognitive intelligence.
연구 동기 및 목표
- 자원 및 언어 격차를 English-language 커뮤니티가 지배하는 foundation 모델에서 해소한다.
- 모델, 도구 및 벤치마크를 통합한 포괄적이고 사용자 중심의 중국어 foundation-model 생태계를 구축한다.
- 중국어 대형 모델 커뮤니티를 발전시키기 위한 오픈 소스 거버넌스와 협업을 제공한다.
제안 방법
- User-Centered Taxonomy (UCT)를 정의하여 사용자 필요를 분류하고 모델 제안에 매핑한다.
- NLU, NLG, NLT 및 멀티모달/도메인/탐색 작업에 걸친 49개의 중국어 모델 카탈로그를 구성하고 오픈 소스화한다(이름 규칙 포함).
- Fengshen Framework를 개발하여 표준 데이터 처리, 모델 인터페이스, 튜토리얼, docker-like 환경 및 업계 표준 API(HuggingFace/Megatron-LM/DeepSpeed 통합)를 결합한다.
- Chinese SuperGLUE와 같은 중국어 대형 벤치마크 및 지식 기반 QA 벤치마크(QAKM) 등 공정하고 미래 지향적인 평가를 가능하게 하는 Fengshenbang Benchmark를 만든다.
- 모델 설계, 선택 기준(파워, 다양성, 사용성) 및 쉽게 발견할 수 있는 명명 체계를 설명한다.
실험 결과
연구 질문
- RQ1종합적이고 표준화되며 사용자 중심의 중국어 foundation-model 생태계가 어떻게 설계되고 평가될 수 있는가?
- RQ2중국어 NLP의 진보와 접근성을 가장 잘 지원하는 모델 분류체계, 명명법, 선택 기준은 무엇인가?
- RQ3도구 및 벤치마크가 연구자와 실무자들이 다양한 자원으로도 공정한 비교와 사용의 용이성을 어떻게 제공할 수 있는가?
주요 결과
- Fengshenbang를 세 부분 생태계로 소개: Fengshenbang Model, Fengshen Framework, Fengshen Benchmark.
- 명확한 명명 규칙과 사용자 중심의 분류 체계로 49개의 오픈 중국어 모델을 출시하고 문서화한다.
- HuggingFace, Megatron-LM, PyTorch-Lightning, DeepSpeed를 통합하는 프레임워크를 만들어 매우 큰 모델(10B 매개변수 이상)의 학습 및 미세 조정을 가능하게 한다.
- 중국어 중심 벤치마크를 개발하고 중국어-SuperGLUE 및 QAKM에 대한 계획을 포함하여 공정한 평가와 진행 상황 추적을 지원한다.
- 사전 학습 Chinese 모델을 선택하고 Fengshen Framework 튜토리얼로 정제한 후 Fengshenbang Benchmarks 또는 맞춤형 작업에서 평가하는 실용적인 3단계 사용 흐름.
- 중국어 오픈 소스 모델 생태계를 형성하기 위한 윤리적 고려사항과 지속적인 커뮤니티 주도 개발을 강조한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.