Skip to main content
QUICK REVIEW

[论文解读] META-pipe - Pipeline Annotation, Analysis and Visualization of Marine Metagenomic Sequence Data

Espen Mikal Robertsen, Tim Kahlke|arXiv (Cornell University)|Apr 14, 2016
Scientific Computing and Data Management参考文献 43被引用 14
一句话总结

META-pipe 是一种可扩展的、与云集成的海洋宏基因组数据分析流水线,结合了分布式计算和存储资源的预处理、组装、分类群分类和功能注释。它利用现有的 Galaxy 框架和超算基础设施,实现高效、交互式的可视化和工作流执行,展示了在大规模海洋宏基因组学项目中高性能和可扩展性的特点。

ABSTRACT

The marine environment is one of the most important sources for microbial biodiversity on the planet. These microbes are drivers for many biogeochemical processes, and their enormous genetic potential is still not fully explored or exploited. Marine metagenomics (DNA shotgun sequencing), not only offers opportunities for studying structure and function of microbial communities, but also identification of novel biocatalysts and bioactive compounds. However, data analysis, management, storage, processing and interpretation are significant challenges in marine metagenomics due to the high diversity in samples and the size of the marine flagship projects. We provide a new pipeline, META-pipe, for marine metagenomics analysis. It offers pre- processing, assembly, taxonomic classification and functional analysis. To reduce the effort to develop and deploy it, we have integrated existing biological analysis frameworks, and compute and storage infrastructure resources. Our current META-pipe web service provides integration with identity provider services, distributed storage, computation on a Supercomputer, Galaxy workflows, and interactive data visualizations. We have evaluated the scalability and performance of the analysis pipeline. Our results demonstrate how to develop and deploy a pipeline on distributed compute and storage resources, and discusses important challenges related to this process.

研究动机与目标

  • 解决管理、处理和分析高微生物多样性大规模海洋宏基因组数据集所面临的挑战。
  • 通过与现有生物信息学框架和分布式基础设施集成,降低部署和维护宏基因组分析流水线的复杂性。
  • 通过 Galaxy 工作流和分布式存储,实现对海洋宏基因组数据的交互式、可扩展且可重现的分析。
  • 通过提供端到端的功能和分类群注释,支持新微生物功能和生物催化剂的发现。

提出的方法

  • 该流水线使用成熟的生物信息学工具,将预处理、序列组装、分类群分类和功能注释整合到单一工作流中。
  • 它利用超算进行高性能计算,并使用分布式存储实现可扩展的数据处理。
  • 系统使用身份提供者服务,实现对基于网页界面的安全访问和身份验证。
  • 嵌入 Galaxy 工作流,以支持可重现且模块化的分析流水线。
  • 实现交互式数据可视化,以支持对分类群和功能谱的实时探索。
  • 该流水线设计具有可扩展性,支持新工具和分析模块的集成。

实验结果

研究问题

  • RQ1如何设计一种可扩展的、分布式的流水线,以应对大规模海洋宏基因组数据集的计算和存储需求?
  • RQ2哪些架构模式能够实现现有生物信息学工具与高性能计算和云存储的高效集成?
  • RQ3该流水线如何支持宏基因组数据的可重现性、安全性以及交互式探索?
  • RQ4在真实世界的海洋宏基因组学工作负载下,该流水线实现了怎样的性能和可扩展性特征?
  • RQ5此类流水线如何促进在海洋环境中发现新型微生物功能和生物催化剂?

主要发现

  • 该流水线成功集成了分布式存储、超算资源和 Galaxy 工作流,实现了可扩展的宏基因组分析。
  • META-pipe 支持对分类群和功能谱的交互式可视化,提升了数据解读能力。
  • 该系统表现出高性能和可扩展性,适用于大规模海洋宏基因组学项目。
  • 使用现有框架和基础设施显著降低了开发开销,并提高了可维护性。
  • 该流水线通过身份提供者集成实现了安全、经过认证的访问,支持协作研究。
  • 该架构具有可扩展性,支持未来新工具和分析模块的集成。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。