Creating Synthetic Datasets for Collaborative Filtering Recommender Systems using Generative Adversarial Networks

03/02/2023
by   Jesús Bobadilla, et al.
0

Research and education in machine learning needs diverse, representative, and open datasets that contain sufficient samples to handle the necessary training, validation, and testing tasks. Currently, the Recommender Systems area includes a large number of subfields in which accuracy and beyond accuracy quality measures are continuously improved. To feed this research variety, it is necessary and convenient to reinforce the existing datasets with synthetic ones. This paper proposes a Generative Adversarial Network (GAN)-based method to generate collaborative filtering datasets in a parameterized way, by selecting their preferred number of users, items, samples, and stochastic variability. This parameterization cannot be made using regular GANs. Our GAN model is fed with dense, short, and continuous embedding representations of items and users, instead of sparse, large, and discrete vectors, to make an accurate and quick learning, compared to the traditional approach based on large and sparse input vectors. The proposed architecture includes a DeepMF model to extract the dense user and item embeddings, as well as a clustering process to convert from the dense GAN generated samples to the discrete and sparse ones, necessary to create each required synthetic dataset. The results of three different source datasets show adequate distributions and expected quality values and evolutions on the generated datasets compared to the source ones. Synthetic datasets and source codes are available to researchers.

READ FULL TEXT

page 5

page 23

research
12/26/2018

Deep Item-based Collaborative Filtering for Sparse Implicit Feedback

Recommender systems are ubiquitous in the domain of e-commerce, used to ...
research
09/12/2021

An Improved Hybrid Recommender System: Integrating Document Context-Based and Behavior-Based Methods

One of the main challenges in recommender systems is data sparsity which...
research
06/05/2020

Providing reliability in Recommender Systems through Bernoulli Matrix Factorization

Recommender Systems are giving increasing importance to the beyond accur...
research
07/18/2018

Trust-Based Collaborative Filtering: Tackling the Cold Start Problem Using Regular Equivalence

User-based Collaborative Filtering (CF) is one of the most popular appro...
research
11/01/2016

The Deep Journey from Content to Collaborative Filtering

In Recommender Systems research, algorithms are often characterized as e...
research
08/14/2017

Collaborative Filtering using Denoising Auto-Encoders for Market Basket Data

Recommender systems (RS) help users navigate large sets of items in the ...
research
01/03/2023

Multidimensional Item Response Theory in the Style of Collaborative Filtering

This paper presents a machine learning approach to multidimensional item...

Please sign up or login with your details

Forgot password? Click here to reset