[Paper Review] A monitoring tool for a GRID operation center
This paper presents EDT-Monitor, a Nagios-based monitoring tool developed for the WorldGRID intercontinental Grid testbed to enable Virtual Organization-centric monitoring through dynamic geographical maps. It enhances Grid operation center capabilities by detecting issues like malfunctioning Resource Brokers, misconfigured sites, and job dispatching failures in real time, improving system reliability and operational responsiveness during large-scale HEP workloads.
WorldGRID is an intercontinental testbed spanning Europe and the US integrating architecturally different Grid implementations based on the Globus toolkit. The WorldGRID testbed has been successfully demonstrated during the WorldGRID demos at SuperComputing 2002 (Baltimore) and IST2002 (Copenhagen) where real HEP application jobs were transparently submitted from US and Europe using "native" mechanisms and run where resources were available, independently of their location. To monitor the behavior and performance of such testbed and spot problems as soon as they arise, DataTAG has developed the EDT-Monitor tool based on the Nagios package that allows for Virtual Organization centric views of the Grid through dynamic geographical maps. The tool has been used to spot several problems during the WorldGRID operations, such as malfunctioning Resource Brokers or Information Servers, sites not correctly configured, job dispatching problems, etc. In this paper we give an overview of the package, its features and scalability solutions and we report on the experience acquired and the benefit that a GRID operation center would gain from such a tool.
Motivation & Objective
- To address the challenge of monitoring complex, geographically distributed Grid infrastructures with heterogeneous components.
- To provide real-time visibility into the health and performance of Grid resources across multiple continents.
- To support Virtual Organization-centric monitoring to simplify operational oversight in large-scale Grid deployments.
- To detect and report infrastructure issues—such as failed brokers or misconfigured sites—proactively during production operations.
- To improve the scalability and maintainability of Grid operation centers through automated, centralized monitoring.
Proposed method
- The EDT-Monitor tool is built on the Nagios open-source monitoring framework to leverage its proven infrastructure for service and host monitoring.
- It integrates with Globus Toolkit-based Grid components, including Resource Brokers and Information Servers, to collect real-time status data.
- The tool generates dynamic geographical maps that visualize the operational status of Grid sites across Europe and the US, enabling spatial correlation of failures.
- It supports Virtual Organization (VO)-centric views, allowing administrators to monitor resources based on VO membership and job submission patterns.
- The system uses event-driven alerts and status aggregation to scale across hundreds of distributed Grid sites and services.
- It enables correlation of monitoring data with actual job execution behavior, such as job dispatching delays or resource unavailability.
Experimental results
Research questions
- RQ1How can a monitoring system effectively support real-time, large-scale Grid operations across geographically distributed, heterogeneous infrastructures?
- RQ2What mechanisms enable Virtual Organization-centric monitoring in a multi-domain Grid environment?
- RQ3How can dynamic geographical visualization improve fault detection and operational response in Grid operation centers?
- RQ4What scalability and reliability features are required for monitoring tools in production Grid testbeds?
- RQ5How can monitoring tools detect and report infrastructure issues such as failed brokers or misconfigured sites before they impact job execution?
Key findings
- EDT-Monitor successfully detected and reported multiple operational issues during WorldGRID demonstrations, including malfunctioning Resource Brokers and misconfigured sites.
- The tool enabled real-time identification of job dispatching problems, reducing mean time to detect failures in the Grid infrastructure.
- Dynamic geographical maps provided intuitive, spatially aware visualization of Grid status, improving situational awareness for operation center staff.
- The integration with Nagios ensured robust, scalable monitoring of hundreds of Grid components with minimal performance overhead.
- The VO-centric monitoring model allowed administrators to isolate and troubleshoot issues based on organizational boundaries, improving operational efficiency.
- The tool demonstrated practical utility during major events like SuperComputing 2002 and IST2002, confirming its effectiveness in real-world Grid operations.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.