Text Mining with R
21 apr. 2022 - 22 apr. 2022
Statistical short course by Leuven Statistics Research Centre
Praktische info:
21 apr. 2022 - 22 apr. 2022
Vandenheuvelinstituut, Dekenstraat 2, 3000 Leuven; 21/04: VHI 02.24 (PC-klas F2); 22/04: VHI 01.23 (PC-klas E1)
Engels
Doelgroep: R users met interesse in natural language processing en text data
Inschrijven?
- Inschrijvingen: tot 22 apr. 2022
- Voorwaarden: Basiskennis van R
- Prijs: KU Leuven studenten €100; KU Leuven medewerkers en andere studenten €160; Non profit/sociale sector €250; Privésector €600
Leertraject
This course is a hands-on course covering the use of text mining tools for the purpose of data analysis. It covers basic text handling, natural language engineering and statistical modelling on top of textual data.
The following items are covered :
- Cleaning of text data, regular expressions
- String distances
- Graphical displays of text data
- Natural language processing: stemming, parts-of-speech tagging, tokenization, lemmatisation
- Sentiment analysis
- Statistical topic detection modelling (latent diriclet allocation)
- Automatic classification using predictive modelling based on text data
- Visualisation of correlations & topics
- Word embeddings
- Document similarities & Text alignment
Lesgever/spreker
Jan Wijffels
Jan Wijffels is the founder of www.bnosac.be - a consultancy company specialised in statistical analysis and data mining. He holds a Master in Commercial Engineering, a MSc in Statistics and a Master in Artificial Intelligence and has been using R for 10 years, developing and deploying R-based solutions for clients in the private sector. He has developed and co-developed the R packages ffbase, ETLUtils, RMOA and RMyrrix.
Gerelateerde opleidingen