On Detecting Hidden Third-Party Web Trackers with a Wide Dependency Chain Graph: A Representation Learning Approach

05/01/2020
by   Amir Hossein Kargaran, et al.
1

Websites use third-party ads and tracking services to deliver targeted ads and collect information about users that visit them. These services put users privacy at risk and that's why users demand to block these services is growing. Most of the blocking solutions rely on crowd-sourced filter lists that are built and maintained manually by a large community of users. In this work, we seek to simplify the update of these filter lists by automatic detection of hidden advertisements. Existing tracker detection approaches generally focus on each individual website's URL patterns, code structure and/or DOM structure of website. Our work differs from existing approaches by combining different websites through a large scale graph connecting all resource requests made over a large set of sites. This graph is thereafter used to train a machine learning model, through graph representation learning to detect ads and tracking resources. As our approach combines different sources of information, it is more robust toward evasion techniques that use obfuscation or change usage patterns. We evaluate our work over the Alexa top-10K websites, and find its accuracy to be 90.9% also it can block new ads and tracking services which would necessitate to be blocked further crowd-sourced existing filter lists. Moreover, the approach followed in this paper sheds light on the ecosystem of third party tracking and advertising

READ FULL TEXT

page 2

page 3

page 4

page 5

page 6

page 7

page 8

page 9

research
06/01/2019

A Longitudinal Analysis of Online Ad-Blocking Blacklists

Websites employ third-party ads and tracking services leveraging cookies...
research
05/22/2018

AdGraph: A Machine Learning Approach to Automatic and Effective Adblocking

Filter lists are widely deployed by adblockers to block ads and other fo...
research
09/12/2023

Cookiescanner: An Automated Tool for Detecting and Evaluating GDPR Consent Notices on Websites

The enforcement of the GDPR led to the widespread adoption of consent no...
research
09/29/2020

A machine learning approach for detecting CNAME cloaking-based tracking on the Web

Various in-browser privacy protection techniques have been designed to p...
research
01/26/2023

ASTrack: Automatic Detection and Removal of Web Tracking Code with Minimal Functionality Loss

Recent advances in web technologies make it more difficult than ever to ...
research
08/07/2023

PURL: Safe and Effective Sanitization of Link Decoration

While privacy-focused browsers have taken steps to block third-party coo...
research
10/20/2020

Grow your Business with Targeted Email List

Email Scraping Services is a very fast and supple #website scraper and e...

Please sign up or login with your details

Forgot password? Click here to reset