Auditing ImageNet: Towards a Model-driven Framework for Annotating Demographic Attributes of Large-Scale Image Datasets

05/03/2019
by   Chris Dulhanty, et al.
0

The ImageNet dataset ushered in a flood of academic and industry interest in leveraging deep learning for computer vision applications. Despite the significant impact of the dataset on the field, there has not been a comprehensive investigation into the demographic attributes of the images contained within this dataset. Such an investigation could lead to new insights on inherent biases deep within the dataset, which is particularly important given it is frequently used to pretrain models for a wide variety of computer vision tasks. In this study, we introduce a model-driven framework for the automatic annotation of apparent age and gender attributes in large-scale image datasets. Using this framework, we conduct a comprehensive demographic audit of the 2012 ImageNet Large Scale Visual Recognition Challenge (ILSVRC) subset of ImageNet and the 'person' hierarchical category of ImageNet by studying the resulting annotations. We find that 41.62 1.71 account for the largest subgroup, at 27.11 apparent demographics of ImageNet are important to identify so the indirect effects of such biases can be better studied. Code and annotations for this work are available at: http://bit.ly/ImageNetDemoAudit

READ FULL TEXT

page 1

page 2

page 3

page 4

research
08/11/2022

A Comprehensive Analysis of AI Biases in DeepFake Detection With Massively Annotated Databases

In recent years, image and video manipulations with DeepFake have become...
research
08/31/2023

FACET: Fairness in Computer Vision Evaluation Benchmark

Computer vision models have known performance disparities across attribu...
research
11/30/2020

Person Perception Biases Exposed: Revisiting the First Impressions Dataset

This work revisits the ChaLearn First Impressions database, annotated fo...
research
06/24/2020

Large image datasets: A pyrrhic win for computer vision?

In this paper we investigate problematic practices and consequences of l...
research
12/22/2019

Analyzing ImageNet with Spectral Relevance Analysis: Towards ImageNet un-Hans'ed

Today's machine learning models for computer vision are typically traine...
research
04/09/2020

PANDORA Talks: Personality and Demographics on Reddit

Personality and demographics are important variables in social sciences,...
research
12/16/2019

Towards Fairer Datasets: Filtering and Balancing the Distribution of the People Subtree in the ImageNet Hierarchy

Computer vision technology is being used by many but remains representat...

Please sign up or login with your details

Forgot password? Click here to reset