Data Wrangling in Python
The handling of data is a recurring task for most scientists. Reading in experimental data, checking its properties, and creating visualisations may become tedious tasks. Hence,increasing the efficiency in this process is beneficial for many scientists. Spreadsheet-based software lacks the ability to properly support this process, due to the lack of automation and repeatability. The usage of a high-level scripting language such as Python is ideal for these tasks.
Practical information:
Want to register?
- Prerequisites: basic programming skills
- Price: determined upon registration
The workshops take place on 25/11/2024 and 02/12/2024, from 09:00 until 18:00.
This course trains students to use Python effectively to do these tasks. The course focuses on data manipulation and cleaning, explorative analysis and visualisation using some important packages such as Pandas, Numpy and Matplotlib. The course is scheduled as a two-day course. On the first day, setting up the programming environment with the required packages using the conda package manager and an introduction of the Jupyter notebook environment are covered. Next, the data analysis package Pandas and Matplotlib is introduced. On the second day, more advanced usage of Pandas for different data cleaning and manipulation tasks is taught. The acquired skills will immediately be brought into practice to handle real-world data sets. Applications include time series handling, categorical data, merging data,...
The course does not cover statistics, data mining, machine learning, or predictive modelling. It aims to provide researchers the means to effectively tackle commonly encountered data handling tasks in order to increase the overall efficiency of the research.
Prerequisites
This course is intended for researchers that have at least basic programming skills. A basic(scientific) programming course that is part of the regular curriculum should suffice. For those who have experience in another programming language (e.g. Matlab, R, ...), following a Python tutorial prior to the course is advised.It is intended for researchers that want to enhance their general data manipulation and analysis skills in Python. The course is NOT intended to be a course on statistics or machine learning.
Teachers / speakers
Research Software Engineer & IT Team lead at Fluves
Experienced Bio-engineer with a strong affinity and passion for IT. I'm at my best when I am bridging the gap between research and software engineering. Seasoned as a researcher and teaching assistant in academia, I switched to computing to maximize my impact by building software, data pipelines and services that convert research and development into production grade applications. I use the best of my abilities to solve real-world environmental problems. Passionate about water, teaching, open science, data visualization and scientific computing.
I am Joris Van den Bossche, an open source python enthusiast. I am core contributor to Pandas, and I am currently working at Voltron Data to advance Arrow-based data tools in Python. I am also a freelance teacher and developer.
I did a PhD at Ghent University and VITO in air quality research (assessing spatial variation in urban air pollution using portable monitors). Afterwards, I was engaged in the CurieuzeNeuzen project, a big citizen science air quality project.
I am a core developer of Pandas, the main data analysis library in Python and have given several tutorials on this topic at international conferences (PyData Paris and EuroScipy) and courses at universities. Further, I am one of the maintainers of GeoPandas, a library to make working with spatial vector data in Python easy. I also contributed to scikit-learn (see e.g. blogpost on the ColumnTransformer), and am now working on Apache Arrow.