The Specialized High-Performance Network on Anton 3

01/20/2022
by   Keun Sup Shim, et al.
0

Molecular dynamics (MD) simulation, a computationally intensive method that provides invaluable insights into the behavior of biomolecules, typically requires large-scale parallelization. Implementation of fast parallel MD simulation demands both high bandwidth and low latency for inter-node communication, but in current semiconductor technology, neither of these properties is scaling as quickly as intra-node computational capacity. This disparity in scaling necessitates architectural innovations to maximize the utilization of computational units. For Anton 3, the latest in a family of highly successful special-purpose supercomputers designed for MD simulations, we thus designed and built a completely new specialized network as part of our ASIC. Tightly integrating this network with specialized computation pipelines enables Anton 3 to perform simulations orders of magnitude faster than any general-purpose supercomputer, and to outperform its predecessor, Anton 2 (the state of the art prior to Anton 3), by an order of magnitude. In this paper, we present the three key features of the network that contribute to the high performance of Anton 3. First, through architectural optimizations, the network achieves very low end-to-end inter-node communication latency for fine-grained messages, allowing for better overlap of computation and communication. Second, novel application-specific compression techniques reduce the size of most messages sent between nodes, thereby increasing effective inter-node bandwidth. Lastly, a new hardware synchronization primitive, called a network fence, supports fast fine-grained synchronization tailored to the data flow within a parallel MD application. These application-driven specializations to the network are critical for Anton 3's MD simulation performance advantage over all other machines.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
02/07/2020

Breaking Band: A Breakdown of High-performance Communication

The critical path of internode communication on large-scale systems is c...
research
01/04/2022

Architectural improvements and technological enhancements for the APEnet+ interconnect system

The APEnet+ board delivers a point-to-point, low-latency, 3D torus netwo...
research
03/26/2018

Homa: A Receiver-Driven Low-Latency Transport Protocol Using Network Priorities

Homa is a new transport protocol for datacenter networks. It provides ex...
research
04/25/2018

Processing Database Joins over a Shared-Nothing System of Multicore Machines

To process a large volume of data, modern data management systems use a ...
research
06/02/2018

Datacenter RPCs can be General and Fast

It is commonly believed that datacenter networking software must sacrifi...
research
10/06/2020

Towards a Scalable and Distributed Infrastructure for Deep Learning Applications

Although recent scaling up approaches to train deep neural networks have...
research
04/26/2022

From Sand to Flour: The Next Leap in Granular Computing with NanoSort

The granularity of distributed computing is limited by communication tim...

Please sign up or login with your details

Forgot password? Click here to reset