Ga verder naar de inhoud

Introduction to big data processing

11 dec. 2024 - 18 dec. 2024

The term "Big Data" refers to data that has characteristics (in terms of Volume, Velocity, Veracity, Variety, ...) that make it difficult to analyze using conventional processing methods (e.g., R or Python or SQL on a single machine). In response, novel programming models and frameworks for dealing with Big Data characteristics have been developed. The objective in this course is to introduce researchers to these programming models, where we focus exclusively on the Volume characteristic of Big Data.

Lees meer & inschrijven ⇗

Praktische info:

11 dec. 2024 - 18 dec. 2024
Online
Engels
Doelgroep: researchers

Inschrijven?

  • Voorwaarden: Hands-on knowledge of Python
  • Prijs: determined upon registration
  • Dates: 11 December 2024 & 18 December 2024, from 13:00 to 17:00.

Lees meer & inschrijven ⇗

georganiseerd door:

Concretely, one way of dealing with the Volume characteristic is to analyze such data in parallel, using many machines organized in a distributed compute cluster. We introduce the architecture of distributed compute clusters (a.k.a. "Warehouse Scale Machines"); how such clusters enable data-parallelism; and the importance of fault-tolerance for such clusters.

We then study how to program on such clusters using the Map/Reduce and Apache Spark programming frameworks.

The course does not cover statistics, data mining, machine learning, or predictive modelling. Neither does it constitute an introduction to programming. It aims to provide researchers the means to effectively analyze data at scale, using modern distributed analytics engines to exploit parallellism.

The course is organized as a two-day course. Each day starts with a theoretical part that discusses the necessary background, and is followed by a hands-on exercise session where participants use the corresponding programming models.

Prerequisites

The exercises will interact with Map/Reduce and Spark through python. A hands-on knowledge of python is therefore required. Since the exercises will be done in the form of jupyter notebooks, familiarity with jupyter notebooks is also required

Lesgever/spreker

I am professor of Computer Science, working in the Data Science Institute at Hasselt University.

My research interests lie in the wide area of data and information management where I focus both on foundational and system aspects.