Proceed to contents

Text Mining with Python

27 Apr 2021 - 28 Apr 2021

Statistical short course by Leuven Statistics Research Centre

Practical information:

27 Apr 2021 - 28 Apr 2021
Vandenheuvelinstituut, Dekenstraat 2, 3000 Leuven; 27/04: VHI 01.23 (PC-room E1); 28/04: VHI 01.23 (PC-room E1)
English
Target audience: Python users interested in natural language processing and text data

Want to register?

  • Register until: 28 Apr 2022
  • Prerequisites: Basic knowledge of Python
  • Price: KU Leuven students €100; KU Leuven staff and other students €160; Non profit/social sector €250; Private sector €600
More info & registration ⇗

Georganiseerd door:

This course is a hands-on course covering the use of text mining tools for the purpose of data analysis. It covers basic text handling, natural language engineering and statistical modelling on top of textual data.

The following items are covered :

  • Encodings, cleaning of text data, regular expressions
  • Language identification
  • String distances
  • Graphical displays of text data
  • Natural language processing: parts-of-speech tagging, tokenization, lemmatisation, keyword extraction, named-entity-recognition
  • Sentiment analysis
  • Statistical topic detection modelling using Gensim
  • Automatic classification using predictive modelling based on text data
  • Word embeddings, document similarities & Text alignment

Teacher / speaker

Jan Wijffels is the founder of www.bnosac.be - a consultancy company specialised in statistical analysis and data mining. He holds a Master in Commercial Engineering, a MSc in Statistics and a Master in Artificial Intelligence and has been using R for 10 years, developing and deploying R-based solutions for clients in the private sector. He has developed and co-developed the R packages ffbase, ETLUtils, RMOA and RMyrrix.