Fast OLAP Query Execution in Main Memory on Large Data in a Cluster

09/15/2017
by   Demian Hespe, et al.
0

Main memory column-stores have proven to be efficient for processing analytical queries. Still, there has been much less work in the context of clusters. Using only a single machine poses several restrictions: Processing power and data volume are bounded to the number of cores and main memory fitting on one tightly coupled system. To enable the processing of larger data sets, switching to a cluster becomes necessary. In this work, we explore techniques for efficient execution of analytical SQL queries on large amounts of data in a parallel database cluster while making maximal use of the available hardware. This includes precompiled query plans for efficient CPU utilization, full parallelization on single nodes and across the cluster, and efficient inter-node communication. We implement all features in a prototype for running a subset of TPC-H benchmark queries. We evaluate our implementation using a 128 node cluster running TPC-H queries with 30 000 gigabyte of uncompressed data.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
07/08/2019

LocationSpark: In-memory Distributed Spatial Query Processing and Optimization

Due to the ubiquity of spatial data applications and the large amounts o...
research
12/16/2021

Predictive Price-Performance Optimization for Serverless Query Processing

We present an efficient, parametric modeling framework for predictive re...
research
05/07/2013

Somoclu: An Efficient Parallel Library for Self-Organizing Maps

Somoclu is a massively parallel tool for training self-organizing maps o...
research
03/27/2022

GPU-Powered Spatial Database Engine for Commodity Hardware: Extended Version

Given the massive growth in the volume of spatial data, there is a great...
research
10/20/2021

Scheduling of Graph Queries: Controlling Intra- and Inter-query Parallelism for a High System Throughput

The vast amounts of data used in social, business or traffic networks, b...
research
04/07/2020

A GPU-friendly Geometric Data Model and Algebra for Spatial Queries: Extended Version

The availability of low cost sensors has led to an unprecedented growth ...
research
03/09/2019

RadegastXDB - Prototype of Native XML Database Management System: Technical Report

A lot of advances in the processing of XML data have been proposed in la...

Please sign up or login with your details

Forgot password? Click here to reset