SoK: Training Machine Learning Models over Multiple Sources with Privacy Preservation

12/06/2020
by   Lushan Song, et al.
0

Nowadays, gathering high-quality training data from multiple data controllers with privacy preservation is a key challenge to train high-quality machine learning models. The potential solutions could dramatically break the barriers among isolated data corpus, and consequently enlarge the range of data available for processing. To this end, both academia researchers and industrial vendors are recently strongly motivated to propose two main-stream folders of solutions: 1) Secure Multi-party Learning (MPL for short); and 2) Federated Learning (FL for short). These two solutions have their advantages and limitations when we evaluate them from privacy preservation, ways of communication, communication overhead, format of data, the accuracy of trained models, and application scenarios. Motivated to demonstrate the research progress and discuss the insights on the future directions, we thoroughly investigate these protocols and frameworks of both MPL and FL. At first, we define the problem of training machine learning models over multiple data sources with privacy-preserving (TMMPP for short). Then, we compare the recent studies of TMMPP from the aspects of the technical routes, parties supported, data partitioning, threat model, and supported machine learning models, to show the advantages and limitations. Next, we introduce the state-of-the-art platforms which support online training over multiple data sources. Finally, we discuss the potential directions to resolve the problem of TMMPP.

READ FULL TEXT
research
03/31/2022

Privacy-Preserving Aggregation in Federated Learning: A Survey

Over the recent years, with the increasing adoption of Federated Learnin...
research
11/10/2020

Privacy Preservation in Federated Learning: Insights from the GDPR Perspective

Along with the blooming of AI and Machine Learning-based applications an...
research
02/03/2021

Federated Learning on Non-IID Data Silos: An Experimental Study

Machine learning services have been emerging in many data-intensive appl...
research
01/04/2021

Fusion of Federated Learning and Industrial Internet of Things: A Survey

Industrial Internet of Things (IIoT) lays a new paradigm for the concept...
research
04/17/2023

Crossing Roads of Federated Learning and Smart Grids: Overview, Challenges, and Perspectives

Consumer's privacy is a main concern in Smart Grids (SGs) due to the sen...
research
10/28/2020

Online feature selection for rapid, low-overhead learning in networked systems

Data-driven functions for operation and management often require measure...
research
08/14/2023

Machine Unlearning: Solutions and Challenges

Machine learning models may inadvertently memorize sensitive, unauthoriz...

Please sign up or login with your details

Forgot password? Click here to reset