Evolution of a Web-Scale Near Duplicate Image Detection System

09/18/2022
by   Andrey Gusev, et al.
0

Detecting near duplicate images is fundamental to the content ecosystem of photo sharing web applications. However, such a task is challenging when involving a web-scale image corpus containing billions of images. In this paper, we present an efficient system for detecting near duplicate images across 8 billion images. Our system consists of three stages: candidate generation, candidate selection, and clustering. We also demonstrate that this system can be used to greatly improve the quality of recommendations and search results across a number of real-world applications. In addition, we include the evolution of the system over the course of six years, bringing out experiences and lessons on how new systems are designed to accommodate organic content growth as well as the latest technology. Finally, we are releasing a human-labeled dataset of  53,000 pairs of images introduced in this paper.

READ FULL TEXT

page 1

page 2

page 6

research
11/21/2021

The Impact of Main Content Extraction on Near-Duplicate Detection

Commercial web search engines employ near-duplicate detection to ensure ...
research
07/31/2015

SnowWatch: Snow Monitoring through Acquisition and Analysis of User-Generated Content

We present a system for complementing snow phenomena monitoring with vir...
research
03/24/2021

Finnish Paraphrase Corpus

In this paper, we introduce the first fully manually annotated paraphras...
research
12/16/2021

Use Image Clustering to Facilitate Technology Assisted Review

During the past decade breakthroughs in GPU hardware and deep neural net...
research
04/12/2018

Optimizing Query Evaluations using Reinforcement Learning for Web Search

In web search, typically a candidate generation step selects a small set...
research
02/12/2018

Image Retargetability

Real-world applications could benefit from the ability to automatically ...
research
08/23/2023

DarkDiff: Explainable web page similarity of TOR onion sites

In large-scale data analysis, near-duplicates are often a problem. For e...

Please sign up or login with your details

Forgot password? Click here to reset