One Backward from Ten Forward, Subsampling for Large-Scale Deep Learning

04/27/2021
by   Chaosheng Dong, et al.
0

Deep learning models in large-scale machine learning systems are often continuously trained with enormous data from production environments. The sheer volume of streaming training data poses a significant challenge to real-time training subsystems and ad-hoc sampling is the standard practice. Our key insight is that these deployed ML systems continuously perform forward passes on data instances during inference, but ad-hoc sampling does not take advantage of this substantial computational effort. Therefore, we propose to record a constant amount of information per instance from these forward passes. The extra information measurably improves the selection of which data instances should participate in forward and backward passes. A novel optimization framework is proposed to analyze this problem and we provide an efficient approximation algorithm under the framework of Mini-batch gradient descent as a practical solution. We also demonstrate the effectiveness of our framework and algorithm on several large-scale classification and regression tasks, when compared with competitive baselines widely used in industry.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
06/19/2023

AdaSelection: Accelerating Deep Learning Training through Data Subsampling

In this paper, we introduce AdaSelection, an adaptive sub-sampling metho...
research
12/08/2021

A Preamble Based MAC Mechanism in Ad-Hoc Network

In this paper, we propose a preamble based medium access control (P-MAC)...
research
10/19/2022

Deep Learning Based Two-dimensional Speaker Localization With Large Ad-hoc Microphone Arrays

Deep learning based speaker localization has shown its advantage in reve...
research
08/20/2021

Pre-training for Ad-hoc Retrieval: Hyperlink is Also You Need

Designing pre-training objectives that more closely resemble the downstr...
research
02/14/2016

Large-Scale Reasoning with OWL

With the growth of the Semantic Web in size and importance, more and mor...
research
04/30/2019

Test Selection for Deep Learning Systems

Testing of deep learning models is challenging due to the excessive numb...
research
03/29/2023

A New Deep Learning and XAI-Based Algorithm for Features Selection in Genomics

In the field of functional genomics, the analysis of gene expression pro...

Please sign up or login with your details

Forgot password? Click here to reset