Leveraging Schema Labels to Enhance Dataset Search

01/27/2020
by   Zhiyu Chen, et al.
0

A search engine's ability to retrieve desirable datasets is important for data sharing and reuse. Existing dataset search engines typically rely on matching queries to dataset descriptions. However, a user may not have enough prior knowledge to write a query using terms that match with description text.We propose a novel schema label generation model which generates possible schema labels based on dataset table content. We incorporate the generated schema labels into a mixed ranking model which not only considers the relevance between the query and dataset metadata but also the similarity between the query and generated schema labels. To evaluate our method on real-world datasets, we create a new benchmark specifically for the dataset retrieval task. Experiments show that our approach can effectively improve the precision and NDCG scores of the dataset retrieval task compared with baseline methods. We also test on a collection of Wikipedia tables to show that the features generated from schema labels can improve the unsupervised and supervised web table retrieval task as well.

READ FULL TEXT
research
06/08/2017

Content-Based Table Retrieval for Web Queries

Understanding the connections between unstructured text and semi-structu...
research
09/17/2021

Towards Handling Unconstrained User Preferences in Dialogue

A user input to a schema-driven dialogue information navigation system, ...
research
06/12/2020

Google Dataset Search by the Numbers

Scientists, governments, and companies increasingly publish datasets on ...
research
05/05/2021

WTR: A Test Collection for Web Table Retrieval

We describe the development, characteristics and availability of a test ...
research
10/03/2022

Unsupervised Search Algorithm Configuration using Query Performance Prediction

Search engine configuration can be quite difficult for inexpert develope...
research
10/14/2020

Valentine: Evaluating Matching Techniques for Dataset Discovery

Data scientists today search large data lakes to discover and integrate ...
research
08/06/2019

RSATree: Distribution-Aware Data Representation of Large-Scale Tabular Datasets for Flexible Visual Query

Analysts commonly investigate the data distributions derived from statis...

Please sign up or login with your details

Forgot password? Click here to reset