Privacy-preserving Deep Learning based Record Linkage

11/03/2022
by   Thilina Ranbaduge, et al.
0

Deep learning-based linkage of records across different databases is becoming increasingly useful in data integration and mining applications to discover new insights from multiple sources of data. However, due to privacy and confidentiality concerns, organisations often are not willing or allowed to share their sensitive data with any external parties, thus making it challenging to build/train deep learning models for record linkage across different organizations' databases. To overcome this limitation, we propose the first deep learning-based multi-party privacy-preserving record linkage (PPRL) protocol that can be used to link sensitive databases held by multiple different organisations. In our approach, each database owner first trains a local deep learning model, which is then uploaded to a secure environment and securely aggregated to create a global model. The global model is then used by a linkage unit to distinguish unlabelled record pairs as matches and non-matches. We utilise differential privacy to achieve provable privacy protection against re-identification attacks. We evaluate the linkage quality and scalability of our approach using several large real-world databases, showing that it can achieve high linkage quality while providing sufficient privacy protection against existing attacks.

READ FULL TEXT
research
12/12/2022

Privacy-Preserving Record Linkage

Given several databases containing person-specific data held by differen...
research
11/29/2019

Incremental Clustering Techniques for Multi-Party Privacy-Preserving Record Linkage

Privacy-Preserving Record Linkage (PPRL) supports the integration of sen...
research
12/25/2018

Privacy-Preserving Collaborative Deep Learning with Irregular Participants

With large amounts of data collected from massive sensors, mobile users ...
research
04/19/2021

Large Scale Record Linkage in the Presence of Missing Data

Record linkage is aimed at the accurate and efficient identification of ...
research
03/10/2021

NegDL: Privacy-Preserving Deep Learning Based on Negative Database

In the era of big data, deep learning has become an increasingly popular...
research
06/07/2019

Increasing Transparent and Accountable Use of Data by Quantifying the Actual Privacy Risk in Interactive Record Linkage

Record linkage refers to the task of integrating data from two or more d...
research
12/22/2020

Modeling Deep Learning Based Privacy Attacks on Physical Mail

Mail privacy protection aims to prevent unauthorized access to hidden co...

Please sign up or login with your details

Forgot password? Click here to reset