Efficient Transmission and Reconstruction of Dependent Data Streams via Edge Sampling

08/12/2022
by   Joel Wolfrath, et al.
0

Data stream processing is an increasingly important topic due to the prevalence of smart devices and the demand for real-time analytics. Geo-distributed streaming systems, where cloud-based queries utilize data streams from multiple distributed devices, face challenges since wide-area network (WAN) bandwidth is often scarce or expensive. Edge computing allows us to address these bandwidth costs by utilizing resources close to the devices, e.g. to perform sampling over the incoming data streams, which trades downstream query accuracy to reduce the overall transmission cost. In this paper, we leverage the fact that correlations between data streams may exist across devices located in the same geographical region. Using this insight, we develop a hybrid edge-cloud system which systematically trades off between sampling at the edge and estimation of missing values in the cloud to reduce traffic over the WAN. We present an optimization framework which computes sample sizes at the edge and systematically bounds the number of samples we can estimate in the cloud given the strength of the correlation between streams. Our evaluation with three real-world datasets shows that compared to existing sampling techniques, our system could provide comparable error rates over multiple aggregate queries while reducing WAN traffic by 27-42

READ FULL TEXT
research
01/04/2020

SurveilEdge: Real-time Video Query based on Collaborative Cloud-Edge Deep Learning

The real-time query of massive surveillance video data plays a fundament...
research
10/04/2022

Sampling Streaming Data with Parallel Vector Quantization – PVQ

Accumulation of corporate data in the cloud has attracted more enterpris...
research
10/31/2021

On multiple IoT data streams processing using LoRaWAN

LoraWAN has turned out to be one of the most successful frameworks in Io...
research
01/21/2023

ScaDLES: Scalable Deep Learning over Streaming data at the Edge

Distributed deep learning (DDL) training systems are designed for cloud ...
research
07/02/2020

S2CE: A Hybrid Cloud and Edge Orchestrator for Mining Exascale Distributed Streams

The explosive increase in volume, velocity, variety, and veracity of dat...
research
04/19/2022

Decentralized Control of Distributed Cloud Networks with Generalized Network Flows

Emerging distributed cloud architectures, e.g., fog and mobile edge comp...
research
11/24/2022

Probabilistic Time Series Forecasting for Adaptive Monitoring in Edge Computing Environments

With increasingly more computation being shifted to the edge of the netw...

Please sign up or login with your details

Forgot password? Click here to reset