Machine Learning Pipelines with Modern Big DataTools for High Energy Physics

09/23/2019
by   Matteo Migliorini, et al.
0

The effective utilization at scale of complex machine learning (ML) techniques to HEP use cases poses several technological challenges, most importantly on the actual implementation of dedicated end-to-end data pipelines. A solution to that issue is presented, which allows training neural network classifiers using solutions from the Big Data ecosystems, integrated with tools, software, and platforms common in the HEP environment. In particular Apache Spark is exploited for data preparation and feature engineering, running the corresponding (Python) code interactively on Jupyter notebooks; key integrations and libraries that make Spark capable of ingesting data stored using ROOT and its EOS/XRootD protocol will be described and discussed. Training of the neural network models, defined by means of Keras API, is performed in a distributed fashion on Spark clusters using BigDL with Analytics Zoo and Tensorflow. The implementation and the results of the distributed training are described in details in this work.

READ FULL TEXT
research
09/23/2019

Machine Learning Pipelines with Modern Big Data Tools for High Energy Physics

The effective utilization at scale of complex machine learning (ML) tech...
research
06/10/2019

Making Classical Machine Learning Pipelines Differentiable: A Neural Translation Approach

Classical Machine Learning (ML) pipelines often comprise of multiple ML ...
research
04/16/2018

BigDL: A Distributed Deep Learning Framework for Big Data

In this paper, we present BigDL, a distributed deep learning framework f...
research
05/04/2019

A Survey of Adaptive Resonance Theory Neural Network Models for Engineering Applications

This survey samples from the ever-growing family of adaptive resonance t...
research
08/11/2018

MARVIN: An Open Machine Learning Corpus and Environment for Automated Machine Learning Primitive Annotation and Execution

In this demo paper, we introduce the DARPA D3M program for automatic mac...
research
10/25/2021

Memory visualization tool for training neural network

Software developed helps world a better place ranging from system softwa...
research
08/28/2020

Coffea – Columnar Object Framework For Effective Analysis

The coffea framework provides a new approach to High-Energy Physics anal...

Please sign up or login with your details

Forgot password? Click here to reset