Proceed to contents
3-day of lectures and hands-on projects

GCND Hackathon: Automatic linguistic annotation and speech recognition of dialects

7 Oct 2024 - 9 Oct 2024
Calling students and researchers in linguistics, digital humanities, social sciences, computer science, natural language processing: we have a rich language data source, and we’d like you to come and work on it! Join us for a dynamic two-part course exploring dialect syntax through hands-on projects!

Practical information:

7 Oct 2024 - 9 Oct 2024
Rozier 44, Ghent
Dutch
Target audience: young researchers (PhD students, postdocs)

Want to register?

  • Register until: 01 Sep 2024
  • Prerequisites: basic knowledge of Dutch
  • Price: free of charge
More info & registration ⇗

Georganiseerd door:

Join us for a dynamic two-part course exploring dialect syntax through hands-on projects. During the first part of the course, participants will learn about automatic speech recognition (ASR), data annotation, and visualization, with a focus on dialect data in lectures taught by Veronique Hoste (UGent) and Hugo van Hamme (KU Leuven).

In the second part, participants will work on collaborative projects using data from the Spoken Corpus of Southern Dutch Dialects (CGND) to enhance ASR or annotation tools, gaining practical experience in dialect research. Potential projects will be proposed by CGND-researchers, but participants will also be encouraged to propose their own projects. Projects may involve developing tools for data analysis, address questions about language variation, or creating works incorporating the range of voices and stories in the dataset.

Participants in the hack will have access to:

The parsed corpus of Southern Dutch Dialects (GCND) is a linguistically annotated corpus based on existing dialect recordings from the 1960s and 1970s: Voices from the past. The corpus, which is still being expanded, currently provides about 500 hours of audio-aligned transcriptions from ca. 550 different locations in two layers, one closer to the dialect and one closer to Standard Dutch. About 50 of those are already part-of-speech tagged, automatically parsed and manually corrected. The corpus is meant to facilitate large-scale research into syntactical particularities of the southern Dutch dialects. The course will be taught in English, but basic knowledge of Dutch will be required to be able to work with the GCND.

Course objectives

  • Develop skills in annotating dialect data sets and create and adapt tools for analyzing dialect data.
  • Gain proficiency in annotation of spoken data focusing on Southern Dutch dialects, automatic speech recognition, and collaborative project development.
  • Acquire interdisciplinary research experience and foster collaborations.

Artificial Intelligence in Business and Industry

22 September 2026

Postgraduate - Kortrijk - KU Leuven, PUC - KU Leuven Continue, VAIA

AI Legal

24 September 2026

Masterclass - Brussels - Agoria