Fairness and Bias in Natural Language Processing
How to use Natural Language Processing and set it up with fair algorithms, avoiding a repetition of (historical) prejudices?
Practical information:
Want to register?
- Register until: 15 Mar 2022
- Prerequisites: good knowledge of machine learning
- Price: €140 for professionals & €70 for researches of Flemish universities
Natural Language Processing, or NLP, is rapidly gaining popularity. It appears more and more in everyday life, most notably in the chatbots that assist you on online shops. Meanwhile, developers continue to struggle with the development of NLP applications because the algorithms learn from existing, historical texts and decisions which include mistakes. So how can they make the systems ‘fair’ if they don’t learn from fair input? And how can they ensure that analyses are neutral and unbiased if their source input consists of prejudiced content? In two half days Prof. Tim Van de Cruys (Dep. Linguistics) and Pieter Delobelle (Dep. Computer Science) from KU Leuven will teach you the most important techniques to recognize and avoid bias in NLP.
Programme
Day 1: Introduction to NLP
Tim Van de Cruys, KU Leuven
Introduction to NLP (14h-16h)
- Different paradigms for NLP: symbolic, statistical, neural
- NLP applications
- Examples of bias in NLP applications
Neural architectures
- Word embeddings
- Continuous bag of words
- Convolutional neural networks
- Recurrent neural networks
- Transformer architectures
- Contextual representations and transfer learning
Practical session: word embeddings (16.30h-17.30h)
- Training word embeddings
- Analogy computations
- Gender bias in word embeddings
- Mitigating gender bias
Day 2: Fairness and Bias in NLP
Pieter Delobelle, KU Leuven
Introduction to bias (14h-16h)
- A classification of bias
- Model interpretability
- Measuring fairness
- Debiasing methods
- Transparent machine learning, model cards
Intrinsic and extrinsic measures of fairness
- History
- General fairness and definitions: stereotyping, protected groups, etc…
- Measuring fairness in NLP
- Evaluations in language models (focus because SOTA)
- Case study on evaluating fairness in RobBERT
- Differences with English
- WEAT/SEAT and PCA-based measures
- Issues with evaluations (“bad seeds”, “nordic salmon”, etc…)
Mitigating stereotypes in language models
- Overview of different methods
- Retraining, adapters, projections, …
- Limitations
Practical session: transformer models (16.30h-17.30h)
- Masked word prediction with BERT
- Biased prediction in BERT
- Finetuning a transformer model for classification
- Biased classification
Register
Registration is only complete when the fee has been paid. We will send you the invoice once your registration has been processed.
Cancellation policy: Cancellation free of charge (10% administrative cost) will be possible till the 6th of March. After that date cancellation free of charge will only be possible with a valid reason. If no valid reason presented no reimboursement will apply.
Teachers / speakers
Tim Van de Cruys is associate professor at the Linguistics Department of the Faculty of Arts, KU Leuven. He has previously worked as a CNRS researcher at the IRIT computer science institute in Toulouse. His research field is computational linguistics. He investigates the automatic extraction of semantics from text, focusing mainly on methods of distributional similarity. He has published on factorization methods, tensor algebra, and neural architectures for language processing.
Pieter Delobelle is currently an AI engineer at Aleph Alpha focussing on inference, alignment and fairness of large language models. Previously, he was a postdoctoral researcher at KU Leuven with a specialization in bias and fairness in large language models and he also developed the state-of-the-art Dutch language model RobBERT. He obtained a Masters in Engineering Technology from KU Leuven in 2018 at the Ghent Technology Campus, Belgium. Subsequently, he obtained an Advanced Masters in Artificial Intelligence from KU Leuven, and he stayed on for a Ph.D. in Computer Science under Professor Bettina Berendt and Professor Luc De Raedt, which he started in 2019 and defended in 2023, titled 'Towards fairer foundation models'. His current research on bias and fairness in large language models led to research visits at Weizenbaum Institute and Bocconi University, as well as an internship at Apple Inc.