Skip to main content
QUICK REVIEW

[Paper Review] HERMES: A Hierarchical Broadcast-Based Silicon Photonic Interconnect for Scalable Many-Core Systems

Moustafa Mohamed, Zheng Li|arXiv (Cornell University)|Jan 19, 2014
Photonic and Optical Devices47 references3 citations
TL;DR

HERMES proposes a hierarchical silicon photonic interconnect that combines a low-latency, power-scalable broadcast network with circuit-switched point-to-point links for high-throughput communication in thousand-core systems. By leveraging adiabatic couplers and a 2-ary folded butterfly topology, it achieves linear power scaling (96% efficiency) and reduces global traffic via bloom filter-based filtering and greedy workload migration, outperforming prior designs in scalability, latency, and energy efficiency.

ABSTRACT

Optical interconnection networks, as enabled by recent advances in silicon photonic device and fabrication technology, have the potential to address on-chip and off-chip communication bottlenecks in many-core systems. Although several designs have shown superior power efficiency and performance compared to electrical alternatives, these networks will not scale to the thousands of cores required in the future. In this paper, we introduce Hermes, a hybrid network composed of an optimized broadcast for power-efficient low-latency global-scale coordination and circuit-switch sub-networks for high-throughput data delivery. This network will scale for use in thousand core chip systems. At the physical level, SoI-based adiabatic coupler has been designed to provide low-loss and compact optical power splitting. Based on the adiabatic coupler, a topology based on 2-ary folded butterfly is designed to provide linear power division in a thousand core layout with minimal cross-overs. To address the network agility and provide for efficient use of optical bandwidth, a flow control and routing mechanism is introduced to dynamically allocate bandwidth and provide fairness usage of network resources. At the system level, bloom filter-based filtering for localization of communication are designed for reducing global traffic. In addition, a novel greedy-based data and workload migration are leveraged to increase the locality of communication in a NUCA (non-uniform cache access) architecture. First order analytic evaluation results have indicated that Hermes is scalable to at least 1024 cores and offers significant performance improvement and power savings over prior silicon photonic designs.

Motivation & Objective

  • Address the scalability limitations of existing silicon photonic interconnects in large-scale many-core systems.
  • Overcome the power inefficiency and latency bottlenecks of electrical and conventional optical interconnects in systems with 1,000+ cores.
  • Achieve linear power scalability and low latency for global broadcast while maintaining high bandwidth for local data transfers.
  • Minimize global communication traffic through communication locality enhancement via data and workload migration.
  • Design a system-level architecture that supports efficient, fair, and dynamic bandwidth allocation across optical network resources.

Proposed method

  • Design a novel SoI-based adiabatic coupler for low-loss, compact optical power splitting in the broadcast network.
  • Implement a 2-ary folded butterfly topology to enable linear power division with minimal crosstalk and cross-overs in a 1,024-core layout.
  • Integrate a circuit-switched optical network for high-throughput, low-latency point-to-point communication of long messages.
  • Develop a flow control and routing mechanism to dynamically allocate bandwidth and ensure fair resource usage across the network.
  • Apply bloom filter-based filtering at the global-local interface to localize communication and reduce global traffic.
  • Introduce a greedy-based data and workload migration technique to enhance communication locality in a NUCA (non-uniform cache access) architecture.

Experimental results

Research questions

  • RQ1Can a broadcast-based optical interconnect scale efficiently to 1,024 cores while maintaining low power and latency?
  • RQ2How can optical interconnects achieve linear power scalability without incurring exponential power growth?
  • RQ3To what extent can communication locality be improved through data and workload migration in a many-core system?
  • RQ4How does the hierarchical design reduce global communication overhead and improve network agility?
  • RQ5What is the performance and power efficiency trade-off of the proposed hybrid optical-electrical network compared to state-of-the-art designs?

Key findings

  • HERMES achieves linear power scalability with 96% optical power division efficiency, significantly outperforming bus-based and crossbar designs that scale poorly with core count.
  • The network scales to at least 1,024 cores with a power overhead of O(√N), enabling efficient operation in large-scale systems.
  • Latency is minimized to O(√N) through hierarchical design and optimized routing, approaching the performance of ideal low-latency networks like Iris while remaining scalable.
  • Bandwidth scalability follows an O(1) trend under power constraints, matching or exceeding that of high-bandwidth designs like bus and crossbar networks.
  • Bloom filter-based filtering reduces global traffic by localizing communication, improving network efficiency and reducing contention.
  • Greedy-based data and workload migration significantly enhance communication locality, reducing reliance on global broadcast and lowering overall latency and energy consumption.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.