Introduction to big data processing
The term "Big Data" refers to data that has characteristics (in terms of Volume, Velocity, Veracity, Variety, ...) that make it difficult to analyze using conventional processing methods (e.g., R or Python or SQL on a single machine). In response, novel programming models and frameworks for dealing with Big Data characteristics have been developed. The objective in this course is to introduce researchers to these programming models, where we focus exclusively on the Volume characteristic of Big Data.
Practical information:
Want to register?
- Prerequisites: hands-on knowledge of Python
- Price: determined upon registration
Dates: 11 December 2024 & 18 December 2024, from 13:00 to 17:00.
Concretely, one way of dealing with the Volume characteristic is to analyze such data in parallel, using many machines organized in a distributed compute cluster. We introduce the architecture of distributed compute clusters (a.k.a. "Warehouse Scale Machines"); how such clusters enable data-parallelism; and the importance of fault-tolerance for such clusters.
We then study how to program on such clusters using the Map/Reduce and Apache Spark programming frameworks.
The course does not cover statistics, data mining, machine learning, or predictive modelling. Neither does it constitute an introduction to programming. It aims to provide researchers the means to effectively analyze data at scale, using modern distributed analytics engines to exploit parallellism.
The course is organized as a two-day course. Each day starts with a theoretical part that discusses the necessary background, and is followed by a hands-on exercise session where participants use the corresponding programming models.
Prerequisites
The exercises will interact with Map/Reduce and Spark through python. A hands-on knowledge of python is therefore required. Since the exercises will be done in the form of jupyter notebooks, familiarity with jupyter notebooks is also required
Teacher / speaker
I am professor of Computer Science, working in the Data Science Institute at Hasselt University.
My research interests lie in the wide area of data and information management where I focus both on foundational and system aspects.