Ga verder naar de inhoud

Commands for Self-Driving Cars & Combining Vision and Language

8 apr 2022 14:00 - 15:00

PSI Seminar

Lees meer & inschrijven ⇗

Praktische info:

8 apr 2022 14:00 - 15:00
KU Leuven ESAT, Aula R (00.54)
Engels
Doelgroep: Iedereen met interesse in AI

Inschrijven?

  • Inschrijvingen: tot 08 apr 2022
  • Prijs: Gratis
Lees meer & inschrijven ⇗

georganiseerd door:

Unsupervised Vision-Language Grammar Induction with Shared Structure Modeling

We introduce a new task, unsupervised vision-language (VL) grammar induction. Given an image-caption pair, the goal is to extract a shared hierarchical structure for both image and language simultaneously. We argue that such structured output, grounded in both modalities, is a clear step towards the high-level understanding of multimodal information. Besides challenges existing in conventional visually grounded grammar induction tasks, VL grammar induction requires a model to capture contextual semantics and perform a fine-grained alignment. To address these challenges, we propose a novel method, CLIORA, which constructs a shared vision-language constituency tree structure with context-dependent semantics for all possible phrases in different levels of the tree. It computes a matching score between each constituent and image region, trained via contrastive learning. It integrates two levels of fusion, namely at feature-level and at score-level, so as to allow fine-grained alignment. We introduce a new evaluation metric for VL grammar induction, CCRA, and show a 3.3% improvement over a strong baseline on Flickr30k Entities. We also evaluate our model via two derived tasks, i.e., language grammar induction and phrase grounding, and improve over the state-of-the-art for both.

Authors: Bo Wan (Speaker), Wenjuan Han, Zilong Zheng, Tinne Tuytelaars

Predicting Physical World Destinations for Commands Given to Self-Driving Cars

In recent years, we have seen significant steps taken in the development of self-driving cars. Multiple companies are starting to roll out impressive systems that work in a variety of settings. These systems can sometimes give the impression that full self-driving is just around the corner and that we would soon build cars without even a steering wheel. The increase in the level of autonomy and control given to an AI provides an opportunity for new modes of human-vehicle interaction. However, surveys have shown that giving more control to an AI in self-driving cars is accompanied by a degree of uneasiness by passengers. In an attempt to alleviate this issue, recent works have taken a natural language-oriented approach by allowing the passenger to give commands that refer to specific objects in the visual scene. Nevertheless, this is only half the task as the car should also understand the physical destination of the command, which is what we focus on in this paper. We propose an extension in which we annotate the 3D destination that the car needs to reach after executing the given command and evaluate multiple different baselines on predicting this destination location. Additionally, we introduce a model that outperforms the prior works adapted for this particular setting.

Authors: Dusan Grujicic (Speaker), Thierry Deruyttere, Marie-Francine Moens, Matthew Blaschko

Gerelateerde opleidingen

HPC Computer Architectures for AI and Dedicated Applications

20 juli 2026

Zomerschool - Barcelona - Barcelona Supercomputing Center

European Agentic AI Bootcamp

20 juli 2026

Training - Online - HLRS, AI:AT