Kafka-ML: connecting the data stream with ML/AI frameworks

06/07/2020
by   Cristian Martín, et al.
0

Machine Learning (ML) and Artificial Intelligence (AI) have a dependency on data sources to train, improve and make predictions through their algorithms. With the digital revolution and current paradigms like the Internet of Things, this information is turning from static data into continuous data streams. However, most of the ML/AI frameworks used nowadays are not fully prepared for this revolution. In this paper, we proposed Kafka-ML, an open-source framework that enables the management of TensorFlow ML/AI pipelines through data streams (Apache Kafka). Kafka-ML provides an accessible and user-friendly Web User Interface where users can easily define ML models, to then train, evaluate and deploy them for inference. Kafka-ML itself and its deployed components are fully managed through containerization technologies, which ensure its portability and easy distribution and other features such as fault-tolerance and high availability. Finally, a novel approach has been introduced to manage and reuse data streams, which may lead to the (no) utilization of data storage and file systems.

READ FULL TEXT
research
09/19/2023

AI/ML for Beam Management in 5G-Advanced

In beamformed wireless cellular systems such as 5G New Radio (NR) networ...
research
11/20/2020

AI Governance for Businesses

Artificial Intelligence (AI) governance regulates the exercise of author...
research
11/24/2018

MLModelScope: Evaluate and Measure ML Models within AI Pipelines

The current landscape of Machine Learning (ML) and Deep Learning (DL) is...
research
10/04/2022

Enabling Serverless Deployment of Large-Scale AI Workloads

We propose a set of optimization techniques for transforming a generic A...
research
09/06/2018

Propheticus: Generalizable Machine Learning Framework

Due to recent technological developments, Machine Learning (ML), a subfi...
research
03/21/2019

Towards Standardization of Data Licenses: The Montreal Data License

This paper provides a taxonomy for the licensing of data in the fields o...
research
12/15/2022

A Data Source Dependency Analysis Framework for Large Scale Data Science Projects

Dependency hell is a well-known pain point in the development of large s...

Please sign up or login with your details

Forgot password? Click here to reset