Balancing Constraints and Submodularity in Data Subset Selection

04/26/2021
by   Srikumar Ramalingam, et al.
7

Deep learning has yielded extraordinary results in vision and natural language processing, but this achievement comes at a cost. Most deep learning models require enormous resources during training, both in terms of computation and in human labeling effort. In this paper, we show that one can achieve similar accuracy to traditional deep-learning models, while using less training data. Much of the previous work in this area relies on using uncertainty or some form of diversity to select subsets of a larger training set. Submodularity, a discrete analogue of convexity, has been exploited to model diversity in various settings including data subset selection. In contrast to prior methods, we propose a novel diversity driven objective function, and balancing constraints on class labels and decision boundaries using matroids. This allows us to use efficient greedy algorithms with approximation guarantees for subset selection. We outperform baselines on standard image classification datasets such as CIFAR-10, CIFAR-100, and ImageNet. In addition, we also show that the proposed balancing constraints can play a key role in boosting the performance in long-tailed datasets such as CIFAR-100-LT.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
05/28/2018

Learning From Less Data: Diversified Subset Selection and Active Learning in Image Classification Tasks

Supervised machine learning based state-of-the-art computer vision techn...
research
01/29/2019

Semantic Redundancies in Image-Classification Datasets: The 10 Don't Need

Large datasets have been crucial to the success of deep learning models ...
research
01/03/2019

Learning From Less Data: A Unified Data Subset Selection and Active Learning Framework for Computer Vision

Supervised machine learning based state-of-the-art computer vision techn...
research
04/30/2021

Submodular Mutual Information for Targeted Data Subset Selection

With the rapid growth of data, it is becoming increasingly difficult to ...
research
11/20/2022

MEESO: A Multi-objective End-to-End Self-Optimized Approach for Automatically Building Deep Learning Models

Deep learning has been widely used in various applications from differen...
research
01/30/2023

MILO: Model-Agnostic Subset Selection Framework for Efficient Model Training and Tuning

Training deep networks and tuning hyperparameters on large datasets is c...
research
09/06/2019

Mass Personalization of Deep Learning

We discuss training techniques, objectives and metrics toward mass perso...

Please sign up or login with your details

Forgot password? Click here to reset