Learning similarity measures from data

01/15/2020
by   Bjørn Magnus Mathisen, et al.
16

Defining similarity measures is a requirement for some machine learning methods. One such method is case-based reasoning (CBR) where the similarity measure is used to retrieve the stored case or set of cases most similar to the query case. Describing a similarity measure analytically is challenging, even for domain experts working with CBR experts. However, data sets are typically gathered as part of constructing a CBR or machine learning system. These datasets are assumed to contain the features that correctly identify the solution from the problem features, thus they may also contain the knowledge to construct or learn such a similarity measure. The main motivation for this work is to automate the construction of similarity measures using machine learning, while keeping training time as low as possible. Our objective is to investigate how to apply machine learning to effectively learn a similarity measure. Such a learned similarity measure could be used for CBR systems, but also for clustering data in semi-supervised learning, or one-shot learning tasks. Recent work has advanced towards this goal, relies on either very long training times or manually modeling parts of the similarity measure. We created a framework to help us analyze current methods for learning similarity measures. This analysis resulted in two novel similarity measure designs. One design using a pre-trained classifier as basis for a similarity measure. The second design uses as little modeling as possible while learning the similarity measure from data and keeping training time low. Both similarity measures were evaluated on 14 different datasets. The evaluation shows that using a classifier as basis for a similarity measure gives state of the art performance. Finally the evaluation shows that our fully data-driven similarity measure design outperforms state of the art methods while keeping training time low.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
03/29/2023

ContraSim – A Similarity Measure Based on Contrastive Learning

Recent work has compared neural network representations via similarity-b...
research
03/28/2022

Comparing in context: Improving cosine similarity measures with a metric tensor

Cosine similarity is a widely used measure of the relatedness of pre-tra...
research
12/27/2019

Efficient Data Analytics on Augmented Similarity Triplets

Many machine learning methods (classification, clustering, etc.) start w...
research
05/21/2019

Similarity Measure Development for Case-Based Reasoning- A Data-driven Approach

In this paper, we demonstrate a data-driven methodology for modelling th...
research
01/07/2022

Generalized quantum similarity learning

The similarity between objects is significant in a broad range of areas....
research
12/14/2021

GEO-BLEU: Similarity Measure for Geospatial Sequences

In recent geospatial research, the importance of modeling large-scale hu...

Please sign up or login with your details

Forgot password? Click here to reset