Explainable and Trustworthy Artificial Intelligence
Artificial Intelligence (AI) has come a long way since its first use and application many decades ago. The use of AI and Machine Learning have seen an immense uptake in the 21st century. The techniques developed in the domain were and are successfully applied to a wide variety of problems, both in academia, private and public industry. As this domain became more and more established in recent years, new challenges arose.
Practical information:
Want to register?
- Prerequisites: a higher education degree in computer science and programming experience with Python or a related programming language
- Price: €1750 (on campus), €1450 (online only)
The lessons are taught from 17h30 till 21h, with a sandwich break.
Artificial Intelligence nowadays are complex and sophisticated algorithms that sometimes make it difficult for the humans to understand and interpret the decisions or suggestions of
the AI system.
Explainable AI puts the following properties on the foreground to deliver trust:
- Gaining trust by explaining for example the characteristics of AI output.
- By explaining an AI technique understanding will increase, allowing to investigate if the technique can be transferred to another domain or problem.
- Informing a user about the workings of an AI model so that there is no misinterpretation.
- Confidence of users can be established by using AI models that are explainable, stable but also robust.
- When explaining AI models issues concerning privacy awareness come into play. Private data should not be exposed by the models.
- It is important that actions can be explained. How have we come to specific outcomes and how could we change them?
- Nowadays, a wide variety of people from different background come into contact AI, it is important that they all understand why the system is behaving in such a manner and offer explanations tailored to their needs.
Target audience
The lessons are intended for anyone who has a good professional familiarity with computer science and who would like to get more insights in techniques that can be applied to achieve explainable and trustworthy Artificial Intelligence. Participants have completed a higher education in computer science or have acquired an equivalent experience.
Participants have programming experience with Python or a related programming language.
The lessons can be followed onsite in the UGain classroom or online live & on demand.
Programme
1. Introduction
In this first lesson, we give a short recap of the basics, followed by the explanation of some general terms that are used in the domain of explainable and trustworthy Artificial Intelligence. This introduction will end with the definition of the challenges within this domain.
- Recap the basics: AI, ML and statistics
- Different types of ML: white-box & black-box
- Interpretability vs Explainability
- Human Uncertainty vs Model Uncertainty
- Challenges
Teachers: Femke Ongenae & Sofie Van Hoecke
Date: 30 September 2024
2. White box models
While black box models offer higher accuracy, white box models are easier to explain and to interpret, unfortunately this leads to a lesser predictive capacity. In the area of white box models, several different approaches will be highlighted:
- Linear Regression
- Generalized Additive Models (GAMs)
- Decision Trees
- Rule-based systems and fuzzy logic
Teacher: Daniel Peralta Cámara
Date: 7 October 2024
3. Interpretability & Explainability
Machine learning systems build models that learn to automate complex tasks by learning from examples. How to get insights into how these models work depends on the type of algorithm used. Getting insights into how our models work can be done by looking at how the model works in general (interpretability), versus how a specific prediction of the model was computed (explainability). Additional hypothetical “What if” questions can be asked to allow for counterfactual reasoning, adding to the toolkit of explainability methods.
- Interpretability versus explainability
- Counterfactuals
- Model distillation
- Dependency plots
- Saliency maps
Teacher: Yvan Saeys
Date: 14 October 2024
4. Online & Transfer Learning
Training machine learning systems can be done before use, i.e. when training it on a stack of pictures first and asking it to make sense of new pictures later. However, it can also be done during use. In the latter scenario the system gets updated whilst it is being used. Sometimes this is necessary because training data is (partially) becoming available after commissioning of the system. Sometimes a system is pretrained on one dataset and the developer wants to retrain the system in order to solve another but related problem, i.e. using a machine vision system that is trained to detect cats to now detect dogs. The developer thus leverages the effort put into the training of the earlier system, hence requiring less training time for the novel system. These and other relations between datasets, their application in training models and the problems we solve with those will be explained in this lesson.
- Online learning
- Change detection
- Transfer learning & domain adaptation (foundation models)
Teacher: Matthias Feys
Date: 21 October 2024
5. Hybrid AI
The oldest forms of machine learning entail rule engines that were hand programmed. Newer forms entail algorithms searching for connections themselves. The first are great in explaining how they reach their conclusions. The latter sometimes give superior predictions, being a lot less brittle, but lack that explainability. To get the best of both worlds, these approaches are sometimes combined. Moreover, allowing an expert to guide a machine learning system can sometimes lead to yet again superior predictions.
- Data-driven vs expert-based approaches
- Finding synergies in data-driven and expert-based approaches
- Combining expert knowledge and machine learning
Teachers: Femke Ongenae & Sofie Van Hoecke
Date: 4 November 2024
6. Robustness
Machine learning systems are extremely fragile: small modifications to their input data can cause them to produce wildly incorrect outputs. These modifications are usually imperceptible or seemingly harmless, making them hard to detect. Such "adversarial perturbations" undermine the trustworthiness of our systems and may pose safety issues under certain circumstances. This lesson explains these problems and what you can do about them.
- Adversarial learning
- Evasion attacks & defenses
- Learning theory
Teacher: Jonathan Peck
Date: 18 November 2024
7. Uncertainty
The notion of uncertainty is of major importance in machine learning and constitutes a key element of modern machine learning methodology. In recent years, it has gained attention due to the increasing relevance of machine learning for practical applications, many of which are coming with safety requirements. In this regard, new problems and challenges have been identified by machine learning scholars, many of which call for novel methodological developments. Indeed, while uncertainty has a long tradition in statistics, and many useful concepts for representing and quantifying uncertainty have been developed on the basis of probability theory, recent research has gone beyond traditional approaches and also leverages more general formalisms and uncertainty calculi.
- Aleatoric and epistemic uncertainty
- First-order uncertainty representations (probabilistic models, calibration methods, set-based representations, conformal prediction, etc.)
- Second-order uncertainty representations (Bayesian methods, ensemble methods, density-based methods, etc.)
Teacher: Willem Waegeman
Date: 25 November 2024
8. Bias & Fairness
When training machine learning systems, the training data can be biased, leading to unwanted outcomes. For example, an HR system trained on old hospital personnel data might discriminate against women for doctor positions and against men for nurse positions, due to historical gender biases in these roles. This session will explain these issues, how to avoid them, how to measure bias and what the limitations of avoiding it are. Also advanced bias and fairness issues in large language models, and generative AI more generally, will be covered.
- Various notions of fairness & impossibility theorem
- Different types of bias & methods to debias
- Ethical and legal guidelines
- Learning fair models
- Uncovering model bias
- Biases in Large Language Models and generative AI
Teacher: Tijl De Bie
Date: 2 December 2024
9. Privacy
Sometimes the quality of machine learning system outputs and privacy are at odds and need to be balanced. However, there are techniques that allow the training of machine learning systems on privacy sensitive data, without exposing the data itself. Those techniques and relevant regulation on these practices are explained in this session.
- Pseudonimization
- K-anonymity
- Differential privacy
- Regulation
Teacher: Tijl De Bie
Date: 9 December 2024
10. Use cases
During the last session, some specific use cases in the domain of Explainable and Trustworthy AI will be discussed.
Teachers: Tijl De Bie, Matthias Feys, Femke Ongenae & Sofie Van Hoecke
Date: 16 December 2024
Teachers / speakers
Femke Ongenae
I am a professor of Data analytics for Health and Connected care, at the IDLab research group of Ghent University. I am part of the PREDICT (http://predict.idlab.ugent.be/) and KnoWS (https://knows.idlab.ugent.be/) research teams. These teams perform research into hybrid AI (fusing semantic models and machine learning), explainable AI, expressive semantic reasoning and the incorporation of expert knowledge in data analytics. This research is mainly applied to the domains of predictive healthcare and industry 4.0 in order to realize context-aware and personalized decision support systems.
I work at the IDLab research group as a postdoctoral researcher. My main passion is fostering collaboration with societal and industrial partners in interdisciplinary projects to valorize our research toward truly impactful applications. This is realized through a number of (government sponsored) projects, for which I set up and lead the trajectory from proposal, into project management and valorization and dissemination afterwards. My main research and project focus is on the use of Semantic Web technologies, machine learning & IoT for the delivery of personalized & context-aware services, especially within the eHealth domain. I am also particularly interested in methodologies for capturing domain knowledge from experts & using this knowledge to optimize intelligent agents and the way we interact with them.
Sofie Van Hoecke graduated from the Engineering Department from the Ghent University in 2003. Following up on her studies in computer science, she achieved a PhD in computer science engineering at the Department of Information Technology at the same university on Efficient service management in healthcare. After being a postdoctoral research engineer at the Department of Information Technology, she started as lecturer ICT and ICT research coordinator at the University College West-Flanders. Currently, she is associate professor at Ghent University, IDLab - Data Science Lab.
Her specialties are: multi-sensor and service oriented architectures, novel services, condition monitoring, emotion recognition, machine learning, semantic dashboards, and the fusion of machine learning and semantic technologies, applied in both predictive maintenance and predictive healthcare.
Dr. Daniel Peralta is a post-doctoral researcher at the Department of Applied Mathematics, Computer Science and Statistics of the Faculty of Sciences of Ghent University. He obtained his PhD at the University of Granada (Spain), tackling large-scale fingerprint identification.
His research has focused on machine learning, especially in large-scale scenarios, and has involved several collaborations with industry to apply such techniques on problems ranging from railway maintenance scheduling to compound activity prediction. Within his current position at the VIB, this research is applied on biological data. He currently teaches Big Data Science courses at Ghent University, in the Master of Statistical Data Analysis and the Master in Computer Science.
Yvan Saeys obtained his PhD in computer science from Ghent University. After spending time abroad at the University of the Basque Country (Spain) and the University of Lyon (France) he returned to Belgium and established the Data Mining and Modeling for Biomedicine (DAMBI) group at the VIB Center for Inflammation Research (IRC) in Gent. As of 2015, he is a professor at Ghent University and a principal investigator (group leader) at VIB, where he is heading an interdisciplinary research team of 21 people, consisting of mathematicians, computer scientists, engineers and bioinformaticians. The Saeys lab studies the design and application of novel data mining and machine learning techniques for high-dimensional single-cell omics data, including methods to model cell developmental trajectories and intercellular communication. At the methodological level, the lab studies the robustness and interpretability of machine learning models.
As Q, Matthias focuses on pushing technical innovation at ML6 and guaranteeing that our customers benefit from the latest technical advancements.
This means constantly improving the chapter working, as well as providing clear links to customer projects and making sure that ML6 has the right technological partnerships to maximize customer impact.
Matthias is energized by coaching technical talent, internally, but also the wider ML ecosystem. He is cofounder & organizer of multiple meetups and frequently acts as technical advisor for startups and colleges/universities.
He's also a trusted expert adviser for the Flemish Agency for Innovation and Entrepreneurship to provide insights and review research projects.
He started his career as PhD researcher at Ghent University focusing on the development of novel event extraction algorithms based on deep learning techniques. Having gained expertise in this new and upcoming technology as well as user testing/evaluation, he was convinced of the immediate opportunities in the industry and prematurely stopped his PhD to join Nicolas in the early days of ML6. This passion led Matthias to become Google Developer Experts for GCP and one of the first GDEs for Machine Learning.
Post-doctoral researcher @ Ghent University
I am a post-doctoral researcher at Ghent University, affiliated with the Department of Applied Mathematics, Computer Science and Statistics (TWIST) as well as the Saeys Lab at the VIB Inflammation Research Center. I am also a teaching assistant for the Artificial Intelligence course offered by Ghent University at the Faculty of Sciences, as well as the lecturer for the Mathematics course in Biomedical Sciences.
My main focus of research is the study of adversarial examples. Broadly speaking, adversarial examples are input samples deliberately crafted by a malicious adversary in order to obtain certain specific predictions from a targeted machine learning model. The intent here is usually to cause some form of harm, such as bypassing automated content filters, malware protections or biometric security systems. In my work, I try to devise countermeasures against this form of exploitation.
Aside from research into adversarial examples, I am also interested in issues of fairness in machine learning. In developing and deploying machine learning systems, researchers and practitioners alike are often ignorant of (or deliberately ignore) the disparate impact of their systems on women and minorities. Some of these tools, such as the recommender systems used by Twitter and Facebook, also facilitate the spread of hate and political extremism across the globe. We cannot afford to remain blind to these problems; the field of machine learning must take its social responsibilities seriously.
Willem Waegeman
Willem Waegeman is an associate professor at Ghent University, and a member of the research unit Knowledge-based Systems (KERMIT) of the Department of Data Analysis and Mathematical Modelling. His main interests are machine learning and bioinformatics. Specific interests include multi-target prediction problems, uncertainty quantification, sequence models and deep learning.
Willem Waegeman is an author of more than 100 papers of peer-reviewed journals and conferences, and his work has won several prizes. In recent years he has served on the program committees of leading conferences in his field (ICML, NIPS, ECML/PKDD, AAAI, AISTATS, IJCAI, etc.).
Since 2008 he is lecturing a machine learning course in Ghent. Since 2014 he is also lecturing several introductory math courses in the first bachelor (> 500 students per year). Willem Waegeman is currently supervising eight PhD-students.
Tijl De Bie has been a Full Professor at the University of Ghent since 2015. Before moving to Ghent, he was Assistant and Associate Professor at the University of Bristol (9 years), a postdoc at the KU Leuven (1 years) and the University of Southampton (1 year). He completed his PhD on machine learning and advanced optimization techniques in 2005 at the KU Leuven. During his PhD he also spent a combined total of about 1 year as a visiting research scholar in U.C. Berkeley and U.C. Davis.
He is currently most actively interested in data-driven Artificial Intelligence (AI), and more specifically in the foundations and applications of (exploratory) Data Science. His focus is increasingly on automating Data Science, human-centric AI (interactivity, privacy, explainability, fairness), and structured data such as graphs. He currently holds a grant portfolio of around EUR 4M, including an ERC Consolidator Grant titled “Formalizing Subjective Interestingness in Exploratory Data Mining” (FORSIED), as well as an FWO Odysseus grant titled “Exploring Data: Theoretical Foundations and Applications to Web, multimedia, and Omics Data”.