A Dynamic, Hierarchical Resource Model for Converged Computing

09/08/2021
by   Daniel J. Milroy, et al.
0

Extreme dynamic heterogeneity in high performance computing systems and the convergence of traditional HPC with new simulation, analysis, and data science approaches impose increasingly more complex requirements on resource and job management software (RJMS). However, there is a paucity of RJMS techniques that can solve key technical challenges associated with those new requirements, particularly when they are coupled. In this paper, we propose a novel dynamic and multi-level resource model approach to address three key well-known challenges individually and in combination: i.e., 1) RJMS dynamism to facilitate job and workflow adaptability, 2) integration of specialized external resources (e.g. user-centric cloud bursting), and 3) scheduling cloud orchestration framework tasks. The core idea is to combine a dynamic directed graph resource model with fully hierarchical scheduling to provide a unified solution to all three key challenges. Our empirical and analytical evaluations of the solution using our prototype extension to Fluxion, a production hierarchical graph-based scheduler, suggest that our unified solution can significantly improve flexibility, performance and scalability across all three problems in comparison to limited traditional approaches.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
08/18/2021

ROME: A Multi-Resource Job Scheduling Framework for Exascale HPC Systems

High-performance computing (HPC) is undergoing significant changes. Next...
research
09/20/2021

Job Scheduling in High Performance Computing

The ever-growing processing power of supercomputers in recent decades en...
research
07/26/2018

Jupyter as Common Technology Platform for Interactive HPC Services

The Minnesota Supercomputing Institute has implemented Jupyterhub and th...
research
10/10/2017

Decentralized Resource Discovery and Management for Future Manycore Systems

The next generation of many-core enabled large-scale computing systems r...
research
08/29/2023

Practice of Alibaba Cloud on Elastic Resource Provisioning for Large-scale Microservices Cluster

Cloud-native architecture is becoming increasingly crucial for today's c...
research
05/15/2020

DeepSoCS: A Neural Scheduler for Heterogeneous System-on-Chip Resource Scheduling

In this paper, we present a novel scheduling solution for a class of Sys...
research
04/19/2021

Mapping the Internet: Modelling Entity Interactions in Complex Heterogeneous Networks

Even though machine learning algorithms already play a significant role ...

Please sign up or login with your details

Forgot password? Click here to reset