Multivariate Microaggregation of Set-Valued Data

04/04/2022
by   Malik Imran-Daud, et al.
0

Data controllers manage immense data, and occasionally, it is released publically to help the researchers to conduct their studies. However, this publically shared data may hold personally identifiable information (PII) that can be collected to re-identify a person. Therefore, an effective anonymization mechanism is required to anonymize such data before it is released publically. Microaggregation is one of the Statistical Disclosure Control (SDC) methods that are widely used by many researchers. This method adapts the k-anonymity principle to generate k-indistinguishable records in the same clusters to preserve the privacy of the individuals. However, in these methods, the size of the clusters is fixed (i.e., k records), and the clusters generated through these methods may hold non-homogeneous records. By considering these issues, we propose an adaptive size clustering technique that aggregates homogeneous records in similar clusters, and the size of the clusters is determined after the semantic analysis of the records. To achieve this, we extend the MDAV microaggregation algorithm to semantically analyze the unstructured records by relying on the taxonomic databases (i.e., WordNet), and then aggregating them in homogeneous clusters. Furthermore, we propose a distance measure that determines the extent to which the records differ from each other, and based on this, homogeneous adaptive clusters are constructed. In experiments, we measured the cohesiveness of the clusters in order to gauge the homogeneity of records. In addition, a method is proposed to measure information loss caused by the redaction method. In experiments, the results show that the proposed mechanism outperforms the existing state-of-the-art solutions.

READ FULL TEXT

page 7

page 9

page 10

page 13

page 17

page 18

page 21

page 22

research
07/16/2018

Novel Feature-Based Clustering of Micro-Panel Data (CluMP)

Micro-panel data are collected and analysed in many research and industr...
research
03/02/2021

Network Cluster-Robust Inference

Since network data commonly consists of observations on a single large n...
research
11/18/2022

Asymptotics for The k-means

The k-means is one of the most important unsupervised learning technique...
research
11/22/2017

Identifying user habits through data mining on call data records

In this paper we propose a framework for identifying patterns and regula...
research
07/11/2014

Biclustering Via Sparse Clustering

In many situations it is desirable to identify clusters that differ with...
research
04/13/2013

Identification of relevant subtypes via preweighted sparse clustering

Cluster analysis methods are used to identify homogeneous subgroups in a...
research
08/12/2017

Bayesian Non-Exhaustive Classification for Active Online Name Disambiguation

The name disambiguation task partitions a collection of records pertaini...

Please sign up or login with your details

Forgot password? Click here to reset