Beyond Text: Multimodal LLM Integration
Multimodal Large Language Models (LLMs) are redefining how we interact with machines—enabling systems that understand not just text, but also speech, video, gaze, and gesture. This Highway Session explores the latest breakthroughs in multimodal AI, with a focus on Flemish innovation and real-world applications across media, automotive, and social robotics.
Join us for an afternoon of cutting-edge talks and networking with researchers and industry experts. From foundational models trained on Flemish media archives to in-vehicle conversational agents and socially aware robots, this session showcases how multimodal LLMs are shaping the future of human–machine interaction.
Practical information:
Want to register?
- Register until: 23 Nov 2025
- Price: free
What can you expect of this study day?
This Highway Session offers a compact and insightful afternoon focused on the evolving landscape of multimodal large language models (LLMs). You’ll be introduced to the latest research directions, practical applications, and strategic initiatives that are shaping how AI systems understand and generate across multiple modalities—text, speech, vision, and beyond:
- A clear overview of current developments in multimodal LLMs
- Reflections on key challenges in benchmarking, integration, and deployment
- Examples of how multimodal AI is being applied in real-world contexts
- Opportunities to connect with peers from academia, industry, and the public sector
Multimodal Foundation Models in the VLAM (Vlaamse AI Modellen) Program
By Tanguy Coenen - imec
The VLAM (Vlaamse AI Modellen) program represents a strategic initiative by the Flemish region to develop robust and culturally attuned foundation models that serve as the backbone for AI innovation across media and other societal sectors. Rooted in the rich digital archives of Flemish media houses—including text, audio, and video content—the program aims to enrich open-source international models with uniquely Flemish data, thereby creating AI systems that reflect the region’s linguistic and cultural identity. VLAM is structured around three core horizons: generative text models, audio synthesis, and video generation. These models are not built from scratch but are adapted from existing foundation base models, ensuring cost-efficiency and technological relevance. A key ambition of VLAM is to empower the Flemish media sector with multimodal interfaces that enable more natural, personalized, and context-aware user experiences. These interfaces could include agentic media agents for consumers and tools for automated content generation and curation for media professionals.
How can we perform fine-grained analysis of multi-modal LLMs?: Performance evaluation and perspectives for future improvements.
By Matthew Blaschko - KU Leuven
Multi-modal LLMs bring additional structure into models and reasoning, such as temporal information in audio and video. While text-based LLMs have impressive performance, albeit with some known shortcomings, less benchmarking has been performed in the multi-modal setting. Despite recent advances in multi-modal learning, existing benchmarks often suffer from strong visual bias – where answers can be inferred from visual data alone – and provide only aggregate scores that conflate multiple sources of error. This makes it difficult to determine whether models struggle with visual understanding, audio interpretation, or audio-visual alignment. In this talk, we present recent insights into how to assess multi-modal LLMs. Our detailed analysis of state-of-the-art models reveals specific failure modes and provides targeted insights for improvement.
Multimodal Language Models for Social Interaction
By Ruben Janssens - UGent
Multimodal Large Language Models (LLMs) have radically transformed our ability to design interactive systems. For the first time, computers can engage in unconstrained, unscripted conversations with people that truly make use of all modalities that carry social interaction. This talk will explore how these models can be integrated into systems that can engage in social conversations, applying them to social robots that help people in educational and healthcare applications. We will cover how multimodal language models can enrich such conversations by combining vision and speech, adapting the conversation to users' non-verbal signals and to the environment — and how they can be safely deployed in sensitive situations by using local models.
Ask the Camera, Get the Why: Zero-Shot Tracking and Natural-Language Explanations
By Nikos Deligiannis - VUB
Imagine saying, “track the red van and the two cyclists near the crossing,” and the AI system in your vehicle does it—then tells you why it made those choices. I will present recent results that make this possible without task-specific training. First, ReferGPT introduces zero-shot referring multi-object tracking: we enrich a multimodal LLM with spatial knowledge so it can produce 3D-aware captions over video frames, then match your natural-language query to those captions using CLIP-style semantic encoding. The result is flexible, open-vocabulary tracking that generalizes to new phrases and achieves competitive performance on automotive benchmarks without bespoke retraining. Second, Zero-Shot Natural Language Explanations deliver faithful, concise “why” answers for any visual classifier without curated explanation data: by aligning class names and concepts to the model’s own embedding space via a tiny neural network trained in seconds, we surface the textual concepts most associated with a prediction. I’ll illustrate how “ask it, track it, explain it” reduces annotation effort and speeds integration for various applications, including vehicle and robot perception.
Multi-Modal In-Vehicle Interaction with CaLLM Edge: The Cerence Automotive LLM
By Bart Baeyens - Cerence
We present recent progress towards fully natural interaction with the vehicle, enabled by CaLLM Edge, the Cerence Automotive Large Language Model. Running at the edge, our system integrates speech, vision, and gaze modalities to create a seamless in-cabin experience. This multi-modal setup allows drivers and passengers to naturally reference and converse about elements both inside and outside the vehicle, opening the door to richer, more intuitive human–vehicle interaction.
Program
13:30 – Welcome & registration
14:00 – Opening by Dr. Sabine Demey (Vlaams AI Onderzoeksprogramma) & Prof. Femke De Backere (VAIA - Vlaamse AI Academie)
14:30 – Multimodal Foundation Models in the VLAM (Vlaamse AI Modellen) Program – Tanguy Coenen (imec)
Multimodal Foundation Models in the VLAM Program
15:00 – How can we perform fine-grained analysis of multi-modal LLMs?: Performance evaluation and perspectives for future improvements – Matthew Blaschko (KU Leuven)
15:45 – Multimodal Language Models for Social Interaction – Ruben Janssens (UGent – IDLab)
16:15 – Ask the Camera, Get the Why: Zero-Shot Tracking and Natural-Language Explanations – Nikos Deligiannis (VUB)
16:45 – Multi-Modal In-Vehicle Interaction with CaLLM Edge: The Cerence Automotive LLM – Bart Baeyens (Cerence)
17:15 – Network Reception
Teachers / speakers
Sabine Demey
Sabine Demey is the director of the Flanders AI Research Program. She brings together researchers from 10 research partners in Flanders (universities and research centres with imec as coordinating partner). Together they tackle challenging AI Research Challenges and apply the new AI methods in healthcare, in industry 5.0, for the energy transition, in society. She believes it is important for technological developments such as AI to have a meaningful impact on people, industry and society. Sabine is a computer scientist with a PhD in robotics. She has 20+ years industrial experience in research, product and business development in 3D printing, software for the manufacturing industry and for healthcare.
Femke De Backere
Femke works part-time at VAIA (the Flemish AI Academy) as a project leader. In addition, she is appointed as a senior lecturer at Ghent University in the field of Software Engineering. The focus of her research lies in making software systems adaptive in an intelligent way. Central to this is the awareness that software solutions do not always work for everyone and that there is no single, one-size-fits-all approach. In terms of applications, she mainly focuses on the broad domain of healthcare—for example, the personalization of physical activity stimulation. This also explains the strong link between her research and her assignments within VAIA. The learning pathways for AI in healthcare closely align with her interests and expertise.
Tanguy Coenen
Dr. Tanguy Coenen is Scientific Director at imec, where he leads strategic initiatives at the intersection of data innovation, AI, and public sector transformation. As the Media and Entertainment Domain Lead and DataTech Scientific Lead, Tanguy plays a pivotal role in shaping the VLAM (Vlaamse AI Modellen) program—Flanders’flagship effort to develop culturally grounded foundation models for generative AI.
With a background in digital transformation and open innovation, Tanguy has consistently championed the creation of open-source, open-data, and open-knowledge deliverables that empower public organizations to tackle complex societal challenges. His work spans domains such as urban logistics, personalized healthcare, and energy flexibility, always with a focus on ethical data sharing and stakeholder collaboration.
In the VLAM program, Tanguy orchestrates cross-sectoral partnerships between imec, media organizations, and academic institutions to build multimodal AI models that reflect Flemish linguistic and cultural nuances. He is also actively involved in shaping the governance and legal frameworks that ensure responsible use of media data in model training. Tanguy is a frequent speaker on topics such as data spaces, agentic AI, and the future of multimodal interfaces in media. He is based in Antwerp and collaborates closely with Flemish and European stakeholders to advance sovereign AI capabilities for the region.
Matthew Blaschko
Matthew Blaschko is a professor in the Center for Processing Speech and Images at KU Leuven. His research is on machine learning foundations and applications in medical image analysis. He is director of the KU Leuven ELLIS unit and a fellow in the ELLIS Health program, part of the European Laboratory for Learning and Intelligent Systems, a principle investigator in the Flanders AI program, and a steering committee member for the KU Leuven AI Institute. He has been a Newton International Fellow at the University of Oxford, as well as a faculty member at the University of Paris-Saclay prior to joining KU Leuven. At KU Leuven, he received the “Golden Chalk” (Gouden Krijtje) award for Best Biomedical Technology Professor 2024-2025.
Ruben Janssens
Ruben Janssens is a final-year PhD student in the AI and Robotics team of IDLab at Ghent University. His research focuses on building artificial intelligence for social robots, enabling them to adapt to visual input in conversations, and is applied to educational robots for language learning.
Nikos Deligiannis
Nikos Deligiannis is a Professor with the Department of Electronics and Informatics (ETRO), Vrije Universiteit Brussel (VUB), the holder of the 2024-2025 Francqui Research Professorship on Trustworthy AI at VUB, and a Principal Investigator at imec, Belgium. He received a 2024 ERC Consolidator Grant to conduct research at the intersection between interpretable and explainable AI and multiterminal compression for autonomous vehicles. His current research focuses on interpretable and explainable AI, multidimensional signal processing, computer vision, and distributed/federated AI for automotive and healthcare applications.
He received the Diploma degree in electrical and computer engineering from the University of Patras, Patras, Greece, in 2006, and the Ph.D. degree in engineering sciences from Vrije Universiteit Brussel (VUB), Brussels, Belgium, in 2012. From 2013 to 2015, he was a Senior Researcher with the Department of Electronic and Electrical Engineering, University College London, London, U.K.
Dr. Deligiannis is a member of the IEEE and EURASIP and served as the Chair of the EURASIP Technical Area Committee on Signal and Data Analytics for Machine Learning in 2021-2023. He serves as an Associate Editor for the IEEE Transactions on Image Processing and has been a Guest Editor for special issues at the EURASIP Journal on Advances in Signal Processing and the Signal Processing journal.
Bart Baeyens
Bart Baeyens is a senior R&D leader with 20+ years of experience in AI and speech technologies. He currently leads global Speech Input R&D at Cerence, delivering ASR and NLU solutions in 35 languages for the automotive sector. With an Executive MBA and a strong engineering background, he specializes in transforming AI research into impactful real-world applications.
Highway: connecting industry with state of the art in AI research
The highway sessions of VAIA bring state of the art of research in artificial intelligence towards the Flemish industry. Joins us for these seminars and stay up to date with the research of today.
Related courses
Summer School on Security and Privacy in the Age of AI
Summerschool - Heverlee - KU Leuven, UGent, VUB, Imec