Attempt to Salvage Multi-million Dollars of Ill-conceived HPC System Investment by Creating Academic Cloud Computing Infrastructure. A Tale of Errors and Belated Learning

05/02/2023
by   Marek Michalewicz, et al.
0

In 2015 the Interdisciplinary Centre for Mathematical and Computational Modelling (ICM), University of Warsaw built a modern datacenter and installed three substantial HPC systems as part of a 168 M PLN (36 M Euro) OCEAN project. Some of the systems were ill-conceived, badly architected and for the five years of their life span have brought minimal ROI. This paper reports on a two-year intensive effort to reengineer two of these HPC systems into a hybrid, multi-cloud solution called A-CHOICeM (Akademicka CHmura Obliczeniowa ICM). The intention was to expand the user base of ICM typical HPC system from around 200 to 500 to about 100,000 potential general academic users from all institutes of higher learning in the Warsaw area. The main characteristics of this solution are integration of on-premises ICM Cloud with several public cloud providers, building solution tailored to particular groups of academic users, containerization, integration of special computational paradigms like AI and Quantum Computing. Full process of designing the solution, competitive dialogue with suppliers, and full final specifications for the solution are presented. Several roadblocks, pitfalls and difficulties encountered along the way, including the conservative attitude of "the old school" HPC admins, University bureaucracy, national funding policies and others are presented.

READ FULL TEXT
research
11/02/2020

10 Years Later: Cloud Computing is Closing the Performance Gap

Can cloud computing infrastructures provide HPC-competitive performance ...
research
12/05/2022

Confidential High-Performance Computing in the Public Cloud

High-Performance Computing (HPC) in the public cloud democratizes the su...
research
03/01/2020

HPC as a Service: A naive model

Applications like Big Data, Machine Learning, Deep Learning and even oth...
research
06/26/2020

Self-Scaling Clusters and Reproducible Containers to Enable Scientific Computing

Container technologies such as Docker have become a crucial component of...
research
03/24/2022

Quantum Computing in the Cloud: Analyzing job and machine characteristics

As the popularity of quantum computing continues to grow, quantum machin...
research
11/18/2022

A DPU Solution for Container Overlay Networks

There is an increasing demand to incorporate hybrid environments as part...
research
07/23/2018

From Bare Metal to Virtual: Lessons Learned when a Supercomputing Institute Deploys its First Cloud

As primary provider for research computing services at the University of...

Please sign up or login with your details

Forgot password? Click here to reset