Supercomputing Academy: Deployable Data Analysis & AI Pipelines with HPC
This course provides an introduction to scalable data analytics, focusing on methods and tools for processing and analysing large-scale datasets on modern high-performance computing systems.
Praktische info:
Inschrijven?
- Voorwaarden: Working knowledge of Linux/Unix & basic Python skills
- Prijs: The price ranges from €0 to €1690
Leertraject
Participants will learn practical patterns to build reproducible and scalable analytics and ML pipelines that remain reliable under real-world constraints, including messy data, growing volumes, performance bottlenecks, and execution across heterogeneous compute environments.
For teams preparing themselves to use HPC resources, the course provides a clear path from notebook prototypes to reproducible, job-based workflows. Participants will develop solutions in Jupyter Notebooks for rapid iteration and then operationalize them as script-based jobs suitable for production-style execution on high-performance systems. The hands-on curriculum covers scalable data processing with Apache Spark, performance-aware execution, and portable environments that help teams turn allocated compute time into measurable progress.
Learning outcomes
After completing this course, participants will be able to:
- Build end-to-end analytics and ML workflows from raw data to evaluated models.
- Perform EDA, data profiling, preprocessing, and feature engineering with practical patterns that reduce downstream issues.
- Implement and compare supervised learning approaches (classification/regression) and select metrics that reflect real objectives.
- Avoid common real-world pitfalls (e.g., data leakage, improper splitting, misleading metrics) and document assumptions clearly.
- Use Apache Spark to scale preprocessing and feature engineering workflows.
- Train and run deep learning models using PyTorch, with practical GPU-aware considerations.
- Understand responsible AI principles and key ethical/legal considerations relevant to real-world usage.
- Gain exposure to emerging HPC/AI trends relevant to modern data and ML workloads.
Agenda
- Week 1: Introduction to Data Analysis and AI within HPC
- Week 2: Machine Learning (ML) and Deep Learning (DL) — from fundamentals to first runnable workflows
- Week 3–4: Practical exercises in data pre-processing, feature engineering, and machine learning applications
- Week 5: Emerging HPC/AI trends — agentic AI, containers, and cost/performance scaling
- Seminars are scheduled on Mondays, 16:30-18:00: Sep. 7 (kick-off, Seminar 1), and Sep. 14, 21, 28, and Oct. 5.
- Exam is scheduled for Friday, Oct. 16. You may start the approximately 2-hour exam anytime between 06:00 and 23:00.
Gerelateerde opleidingen