Navigating the Data Lake with Datamaran: Automatically Extracting Structure from Log Datasets

08/29/2017
by   Yihan Gao, et al.
0

Organizations routinely accumulate semi-structured log datasets generated as the output of code; these datasets remain unused and uninterpreted, and occupy wasted space - this phenomenon has been colloquially referred to as "data lake" problem. One approach to leverage these semi-structured datasets is to convert them into a structured relational format, following which they can be analyzed in conjunction with other datasets. We present Datamaran, an tool that extracts structure from semi-structured log datasets with no human supervision. Datamaran automatically identifies field and record endpoints, separates the structured parts from the unstructured noise or formatting, and can tease apart multiple structures from within a dataset, in order to efficiently extract structured relational datasets from semi-structured log datasets, at scale with high accuracy. Compared to other unsupervised log dataset extraction tools developed in prior work, Datamaran does not require the record boundaries to be known beforehand, making it much more applicable to the noisy log files that are ubiquitous in data lakes. In particular, Datamaran can successfully extract structured information from all datasets used in prior work, and can achieve 95 substantial 66 prior work.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
02/14/2022

vue4logs – Automatic Structuring of Heterogeneous Computer System Logs

Computer system log data is commonly used in system monitoring, performa...
research
07/19/2023

Prompting for Automatic Log Template Extraction

Log parsing, the initial and vital stage in automated log analysis, invo...
research
02/04/2022

Performance Evaluation of Structured and Semi-Structured Bioinformatics Tools: A Comparative Study

There is a wide range of available biological databases developed by bio...
research
11/30/2017

Predicting Severe Sepsis Using Text from the Electronic Health Record

Employing a machine learning approach we predict, up to 24 hours prior, ...
research
04/12/2018

CERES: Distantly Supervised Relation Extraction from the Semi-Structured Web

The web contains countless semi-structured websites, which can be a rich...
research
02/14/2022

UniParser: A Unified Log Parser for Heterogeneous Log Data

Logs provide first-hand information for engineers to diagnose failures i...
research
10/08/2021

Unsupervised Cross-Lingual Transfer of Structured Predictors without Source Data

Providing technologies to communities or domains where training data is ...

Please sign up or login with your details

Forgot password? Click here to reset