In-situ data analytics for highly scalable cloud modelling on Cray machines

10/27/2020
by   Nick Brown, et al.
0

MONC is a highly scalable modelling tool for the investigation of atmospheric flows, turbulence and cloud microphysics. Typical simulations produce very large amounts of raw data which must then be analysed for scientific investigation. For performance and scalability reasons this analysis and subsequent writing to disk should be performed in-situ on the data as it is generated however one does not wish to pause the computation whilst analysis is carried out. In this paper we present the analytics approach of MONC, where cores of a node are shared between computation and data analytics. By asynchronously sending their data to an analytics core, the computational cores can run continuously without having to pause for data writing or analysis. We describe our IO server framework and analytics workflow, which is highly asynchronous, along with solutions to challenges that this approach raises and the performance implications of some common configuration choices. The result of this work is a highly scalable analytics approach and we illustrate on up to 32768 computational cores of a Cray XC30 that there is minimal performance impact on the runtime when enabling data analytics in MONC and also investigate the performance and suitability of our approach on the KNL.

READ FULL TEXT
research
01/20/2021

Neural-based Modeling for Performance Tuning of Spark Data Analytics

Cloud data analytics has become an integral part of enterprise business ...
research
09/12/2019

Simple-ML: Towards a Framework for Semantic Data Analytics Workflows

In this paper we present the Simple-ML framework that we develop to supp...
research
09/27/2020

A highly scalable Met Office NERC Cloud model

Large Eddy Simulation is a critical modelling tool for scientists invest...
research
04/24/2022

Taming Hybrid-Cloud Fast and Scalable Graph Analytics at Twitter

We have witnessed a boosted demand for graph analytics at Twitter in rec...
research
12/30/2021

SIM-SITU: A Framework for the Faithful Simulation of in-situ Workflows

The amount of data generated by numerical simulations in various scienti...
research
03/23/2018

GreyCat: Efficient What-If Analytics for Data in Motion at Scale

Over the last few years, data analytics shifted from a descriptive era, ...
research
09/27/2021

A communication efficient distributed learning framework for smart environments

Due to the pervasive diffusion of personal mobile and IoT devices, many ...

Please sign up or login with your details

Forgot password? Click here to reset