Proceed to contents

Keeping things private: Exploring open-source large language models for sensitive text data

15 Apr 2024 10:00 - 16:00

Do you work with sensitive text data and does that make you hesitant about using AI-models such as ChatGPT? Are you unsure about what really happens to your data, and are you concerned about privacy and security? And most importantly: are you looking for a safer alternative? This course, brought to you by TEXTUA and VAIA, explores the use of open-source large language models as alternatives for dealing with sensitive data in a more secure way. But does their performance match that of the likes of ChatGPT? During a lecture and a hands-on workshop, you gain insight in current large language models' potential and implications for data safety, and you tackle case studies that you can afterwards apply to your own data.

Practical information:

15 Apr 2024 10:00 - 16:00
5 hours
City Campus of the University of Antwerp, Building C, room C.002 (Entrance via Prinsstraat 13, Antwerp)
English
Target audience: researchers from academia and industry, working with sensitive text data

Want to register?

  • Register until: 12 Apr 2024
  • Prerequisites: for the lecture there are no prerequisites; for the workshop a laptop and a basic understanding of Python and the command line are required
  • Price: €50-€150 (see Practical information below)
  • Registrations are open!

More info & registration ⇗

Georganiseerd door:

Course content

PART 1 (morning)

INTRODUCTION by prof. dr. Walter Daelemans:

Welcoming, opening of the course, and introduction to the topic by TEXTUA director professor Walter Daelemans.

LECTURE by dr. Enrique Manjavacas:

While ChatGPT and other large language models (LLMs) have revolutionized research in both academia and industry, such models are typically hosted on external servers which may require undesired (or unclear) sharing of sensitive text data. The emergence of open-source LLMs paves the way for a future in which researchers can 'keep things private' by running smaller, local versions of models. This lecture offers contextual understanding of open-source large language models and explores their current potential for the intelligent processing of sensitive text data.

PART 2 (afternoon)

WORKSHOP by dr. Pieter Fivez:

In this hands-on workshop by TEXTUA coordinator dr. Pieter Fivez, you apply different open-source large language models to prototypical cases of sensitive text data, such as medical data and HR data. After the course, you’ll know how to tackle applying these LLMs to your company’s or research’s own text data, and will have a more developed intuition on the current capabilities and limitations of these models.

The workshop works with Python, in Google Colab Notebooks (which will be provided). See below for prerequisites and equipment.

Target audience & prerequisites

Target audience

The main target audience are researchers from both academia and industry, working with sensitive text data (e.g., patient records, HR files) who want to explore open-source large language models. In addition, anyone wanting to learn about open-source large language models is welcome!

Prerequisites

For the lecture in the morning, there are no prerequisites in terms of background or skills, and no equipment is needed.

For the hands-on workshop in the afternoon, a basic understanding of Python and the command line is required. Google Colab notebooks are provided (so no programming from scratch), but participants need to at least understand the given Python code and be able to tweak it. Advanced programming skills, however, are not required. In case of doubt, contact the workshop teacher Pieter Fivez (Pieter.fivez@uantwerpen.be)

For the workshop, make sure to bring a fully charged laptop with Python installed on it (version 3.9 or higher). Access to Google Colab (free or paying) is required, too.

Programme

Monday 15 April 2024: 10.00 – 16.00h

09.30 – 10.00h: Welcome and coffee

10.00 – 12.00h: Introduction by prof. dr. Walter Daelemans and lecture by dr. Enrique Manjavacas

12.00 – 13.00h: Lunch break (lunch not included)

13.00 – 14.30h: Workshop part 1, by dr. Pieter Fivez

14.30 – 14.45h: Coffee break

14.45 – 16.00h: Workshop part 2, by dr. Pieter Fivez

Note that you can register for:

  • the entire day
  • or only the lecture
  • or only the workshop

Practical information

Price:

  • Lecture (morning): €50
  • Workshop (afternoon): €100
  • Full day (lecture + workshop): €150

Included in the price are: course material, coffee breaks.

Limited places available!

About the speakers and organizers

This course is the very first edition of the 'TEXTUA Invites' series, in which TEXTUA (University of Antwerp) invites national and international experts to tackle diverse challenges within the field of text mining. This course is co-organized by the Flanders AI Academy VAIA.

The lecture is given by international NLP expert dr. Enrique Manjavacas, and the workshop is given by text mining expert and TEXTUA coordinator dr. Pieter Fivez. The opening and introduction is given by TEXTUA director prof. dr. Walter Daelemans.

About TEXTUA

TEXTUA is a core facility of the University of Antwerp, directed by prof. dr. Walter Daelemans and coordinated by dr. Pieter Fivez, which provides scalable text mining solutions to researchers from any scientific discipline. It offers a diverse collection of services for a broad range of textual data, including automatically transcribed speech and written text in images. TEXTUA bundles the unique existing expertise in digital text analysis at the University of Antwerp with special emphasis on explainable Artificial Intelligence.

Teachers / speakers

Enrique Manjavacas

Dr. Enrique Manjavacas is a natural language processing (NLP) engineer with extensive experience in the development and application of large language models (LLMs). He holds a PhD degree in computational linguistics from the University of Antwerp on the application of distributed meaning representations to the extraction of text references. Enrique has a broad interest in machine learning, and specifically in the adaptation and finetuning of LLMs to specific domains. Throughout his career, he has worked on topics including LLM finetuning, word sense disambiguation, morphological analysis, authorship verification, and obfuscation or text generation.

Pieter Fivez

Pieter Fivez holds a PhD in Linguistics from the University of Antwerp, focusing on machine learning of semantic representations of biomedical text. He currently works as a postdoctoral researcher at the University of Antwerp, where he coordinates the Antwerp Text Mining Centre (TEXTUA).

Artificial Intelligence in Business and Industry

22 September 2026

Postgraduate - Kortrijk - KU Leuven, PUC - KU Leuven Continue, VAIA

AI Legal

24 September 2026

Masterclass - Brussels - Agoria