Enabling Quality Control for Entity Resolution: A Human and Machine Cooperation Framework

09/30/2017
by   Zhaoqiang Chen, et al.
0

Even though many machine algorithms have been proposed for entity resolution, it remains very challenging to find a solution with quality guarantees. In this paper, we propose a novel HUman and Machine cOoperation (HUMO) framework for entity resolution (ER), which divides an ER workload between the machine and the human. HUMO enables a mechanism for quality control that can flexibly enforce both precision and recall levels. We introduce the optimization problem of HUMO, minimizing human cost given a quality requirement, and then present three optimization approaches: a conservative baseline one purely based on the monotonicity assumption of precision, a more aggressive one based on sampling and a hybrid one that can take advantage of the strengths of both previous approaches. Finally, we demonstrate by extensive experiments on real and synthetic datasets that HUMO can achieve high-quality results with reasonable return on investment (ROI) in terms of human cost, and it performs considerably better than the state-of-the-art alternatives in quality control.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
09/30/2017

Enabling Quality Control for Entity Resolution: A Human and Machine Cooperative Framework

Even though many machine algorithms have been proposed for entity resolu...
research
03/15/2018

r-HUMO: A Risk-Aware Human-Machine Cooperation Framework for Entity Resolution with Quality Guarantees

Even though many approaches have been proposed for entity resolution (ER...
research
03/15/2018

i-HUMO: An Interactive Human and Machine Cooperation Framework for Entity Resolution with Quality Guarantees

Even though many approaches have been proposed for entity resolution (ER...
research
06/29/2023

UMASS_BioNLP at MEDIQA-Chat 2023: Can LLMs generate high-quality synthetic note-oriented doctor-patient conversations?

This paper presents UMASS_BioNLP team participation in the MEDIQA-Chat 2...
research
05/31/2018

Improving Machine-based Entity Resolution with Limited Human Effort: A Risk Perspective

Pure machine-based solutions usually struggle in the challenging classif...
research
09/10/2015

Performance Bounds for Pairwise Entity Resolution

One significant challenge to scaling entity resolution algorithms to mas...

Please sign up or login with your details

Forgot password? Click here to reset