Embeddings for Tabular Data: A Survey

02/23/2023
by   Rajat Singh, et al.
0

Tabular data comprising rows (samples) with the same set of columns (attributes, is one of the most widely used data-type among various industries, including financial services, health care, research, retail, and logistics, to name a few. Tables are becoming the natural way of storing data among various industries and academia. The data stored in these tables serve as an essential source of information for making various decisions. As computational power and internet connectivity increase, the data stored by these companies grow exponentially, and not only do the databases become vast and challenging to maintain and operate, but the quantity of database tasks also increases. Thus a new line of research work has been started, which applies various learning techniques to support various database tasks for such large and complex tables. In this work, we split the quest of learning on tabular data into two phases: The Classical Learning Phase and The Modern Machine Learning Phase. The classical learning phase consists of the models such as SVMs, linear and logistic regression, and tree-based methods. These models are best suited for small-size tables. However, the number of tasks these models can address is limited to classification and regression. In contrast, the Modern Machine Learning Phase contains models that use deep learning for learning latent space representation of table entities. The objective of this survey is to scrutinize the varied approaches used by practitioners to learn representation for the structured data, and to compare their efficacy.

READ FULL TEXT

page 3

page 5

research
11/14/2019

Sato: Contextual Semantic Type Detection in Tables

Detecting the semantic types of data columns in relational tables is imp...
research
02/01/2020

Web Table Extraction, Retrieval and Augmentation: A Survey

Tables are a powerful and popular tool for organizing and manipulating d...
research
08/23/2022

Data augmentation on graphs for table type classification

Tables are widely used in documents because of their compact and structu...
research
08/29/2017

EntiTables: Smart Assistance for Entity-Focused Tables

Tables are among the most powerful and practical tools for organizing an...
research
08/31/2018

The use of Charts, Pivot Tables, and Array Formulas in two Popular Spreadsheet Corpora

The use of spreadsheets in industry is widespread. Companies base decisi...
research
09/22/2020

On the number of contingency tables and the independence heuristic

We obtain sharp asymptotic estimates on the number of n × n contingency ...
research
11/04/2021

Benchmarking Multimodal AutoML for Tabular Data with Text Fields

We consider the use of automated supervised learning systems for data ta...

Please sign up or login with your details

Forgot password? Click here to reset