Addressing Budget Allocation and Revenue Allocation in Data Market Environments Using an Adaptive Sampling Algorithm

06/05/2023
by   Boxin Zhao, et al.
0

High-quality machine learning models are dependent on access to high-quality training data. When the data are not already available, it is tedious and costly to obtain them. Data markets help with identifying valuable training data: model consumers pay to train a model, the market uses that budget to identify data and train the model (the budget allocation problem), and finally the market compensates data providers according to their data contribution (revenue allocation problem). For example, a bank could pay the data market to access data from other financial institutions to train a fraud detection model. Compensating data contributors requires understanding data's contribution to the model; recent efforts to solve this revenue allocation problem based on the Shapley value are inefficient to lead to practical data markets. In this paper, we introduce a new algorithm to solve budget allocation and revenue allocation problems simultaneously in linear time. The new algorithm employs an adaptive sampling process that selects data from those providers who are contributing the most to the model. Better data means that the algorithm accesses those providers more often, and more frequent accesses corresponds to higher compensation. Furthermore, the algorithm can be deployed in both centralized and federated scenarios, boosting its applicability. We provide theoretical guarantees for the algorithm that show the budget is used efficiently and the properties of revenue allocation are similar to Shapley's. Finally, we conduct an empirical evaluation to show the performance of the algorithm in practical scenarios and when compared to other baselines. Overall, we believe that the new algorithm paves the way for the implementation of practical data markets.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
07/11/2022

Dynamic Budget Throttling in Repeated Second-Price Auctions

Throttling is one of the most popular budget control methods in today's ...
research
11/30/2021

Establishing the Price of Privacy in Federated Data Trading

Personal data is becoming one of the most essential resources in today's...
research
06/27/2021

An Incentive Mechanism for Trading Personal Data in Data Markets

With the proliferation of the digital data economy, digital data is cons...
research
04/24/2023

Dynamic generation and attribution of revenues in a video platform

The consumption of online videos on the Internet grows every year, makin...
research
05/21/2020

Markets for Efficient Public Good Allocation with Social Distancing

Public goods are often either over-consumed in the absence of regulatory...
research
11/08/2019

Collaborative Machine Learning Markets with Data-Replication-Robust Payments

We study the problem of collaborative machine learning markets where mul...
research
06/25/2020

Replication-Robust Payoff-Allocation with Applications in Machine Learning Marketplaces

The ever-increasing take-up of machine learning techniques requires ever...

Please sign up or login with your details

Forgot password? Click here to reset