[论文解读] CASTOR status and evolution
CASTOR,CERN为高能物理工作负载开发的分层存储管理器,自1999年起演进以支持海量数据(拍字节级)和高达300 MB/s的高吞吐量传输,支持关键的LHC数据挑战。关键进展包括支持2GB以上文件及通过GridFTP和SRM实现与网格计算的集成,同时正在进行向更具成本效益的存储介质迁移,并针对数据分析工作负载进行性能优化。
In January 1999, CERN began to develop CASTOR ("CERN Advanced STORage manager"). This Hierarchical Storage Manager targetted at HEP applications has been in full production at CERN since May 2001. It now contains more than two Petabyte of data in roughly 9 million files. In 2002, 350 Terabytes of data were stored for COMPASS at 45 MB/s and a Data Challenge was run for ALICE in preparation for the LHC startup in 2007 and sustained a data transfer to tape of 300 MB/s for one week (180 TB). The major functionality improvements were the support for files larger than 2 GB (in collaboration with IN2P3) and the development of Grid interfaces to CASTOR: GridFTP and SRM ("Storage Resource Manager"). An ongoing effort is taking place to copy the existing data from obsolete media like 9940 A to better cost effective offerings. CASTOR has also been deployed at several HEP sites with little effort. In 2003, we plan to continue working on Grid interfaces and to improve performance not only for Central Data Recording but also for Data Analysis applications where thousands of processes possibly access the same hot data. This could imply the selection of another filesystem or the use of replication (hardware or software).
研究动机与目标
- 解决CERN日益增长的高能物理数据量对可扩展、可靠存储基础设施的需求。
- 支持大型实验(如ALICE和COMPASS)的高吞吐量数据摄入与检索。
- 通过与新兴网格计算基础设施无缝集成,实现分布式数据管理。
- 通过将数据从过时介质(如9940 A)迁移至更具成本效益的解决方案,提升存储效率。
- 针对涉及对热数据集并发访问的数据分析工作负载,优化系统性能。
提出的方法
- 部署CASTOR作为分层存储管理器(HSM),实现磁盘与磁带间的数据自动迁移。
- 通过与IN2P3合作,扩展文件系统支持,以处理超过2 GB的文件。
- 使用GridFTP和SRM(存储资源管理器)协议集成网格接口,实现互操作性。
- 开展持续的数据传输基准测试(例如,一周内保持300 MB/s的传输速率),以验证系统在负载下的性能表现。
- 实施数据迁移管道,将数据从旧的9940 A磁带迁移至新型、更具成本效益的存储介质。
- 实现CASTOR在多个高能物理站点的部署,且配置开销极低。
实验结果
研究问题
- RQ1如何高效管理分层存储系统中高能物理实验的拍字节级数据?
- RQ2在大规模高能物理数据工作负载下,向磁带持续传输数据可达到何种性能水平?
- RQ3如何扩展CASTOR以支持超过2 GB的文件?这是现代高能物理应用的关键需求。
- RQ4CASTOR在多大程度上可通过标准协议(如SRM和GridFTP)与网格计算基础设施实现集成?
- RQ5采用何种策略可实现从过时存储介质向更经济高效介质的迁移,同时不中断数据可用性?
主要发现
- 在ALICE数据挑战期间,CASTOR实现了长达一周的300 MB/s持续数据传输速率,证明其已具备支持LHC规模工作负载的能力。
- 截至2003年,系统已存储超过2 PB的数据,涵盖约900万个文件,证实其具备可扩展性。
- 成功实现180 TB数据在一周内以300 MB/s的速率持续传输,验证了其高吞吐能力。
- 通过与IN2P3合作,成功实现对2 GB以上文件的支持,消除了关键的技术障碍。
- 网格接口(GridFTP和SRM)实现了与分布式计算环境的无缝集成,显著提升了互操作性。
- 从过时的9940 A磁带向新型存储介质的持续迁移,提升了成本效益,同时保持了数据完整性和可用性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。