Bonsai: A Generalized Look at Dual Deduplication

02/28/2022
by   Hadi Sehat, et al.
0

Cloud Service Providers (CSPs) offer a vast amount of storage space at competitive prices to cope with the growing demand for digital data storage. Dual deduplication is a recent framework designed to improve data compression on the CSP while keeping clients' data private from the CSP. To achieve this, clients perform lightweight information-theoretic transformations to their data prior to upload. We investigate the effectiveness of dual deduplication, and propose an improvement for the existing state-of-the-art method. We name our proposal Bonsai as it aims at reducing storage fingerprint and improving scalability. In detail, Bonsai achieves (1) significant reduction in client storage, (2) reduction in total required storage (client + CSP), and (3) reducing the deduplication time on the CSP. Our experiments show that Bonsai achieves compression rates of 68% on the cloud and 5% on the client, while allowing the cloud to identify deduplications in a time-efficient manner. We also show that combining our method with universal compressors in the cloud, e.g., Brotli, can yield better overall compression on the data compared to only applying the universal compressor or plain Bonsai. Finally, we show that Bonsai and its variants provide sufficient privacy against an honest-but-curious CPS that knows the distribution of the Clients' original data.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
07/22/2020

Yggdrasil: Privacy-aware Dual Deduplication in Multi Client Settings

This paper proposes Yggdrasil, a protocol for privacy-aware dual data de...
research
12/12/2017

Keyword-Based Delegable Proofs of Storage

Cloud users (clients) with limited storage capacity at their end can out...
research
04/14/2019

Secure Consistency Verification for Untrusted Cloud Storage by Public Blockchains

This work presents ContractChecker, a Blockchain-based security protocol...
research
08/08/2021

Data Analysis: Communicating with Offshore Vendors using Instant Messaging Services

The purpose of this study is to find whether the choice of correct analy...
research
02/07/2022

Learning under Storage and Privacy Constraints

Storage-efficient privacy-guaranteed learning is crucial due to enormous...
research
09/19/2020

Hierarchical Coding for Cloud Storage: Topology-Adaptivity, Scalability, and Flexibility

In order to accommodate the ever-growing data from various, possibly ind...
research
10/24/2020

Differentiate Quality of Experience Scheduling for Deep Learning Applications with Docker Containers in the Cloud

With the prevalence of big-data-driven applications, such as face recogn...

Please sign up or login with your details

Forgot password? Click here to reset