Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
Please note that all times are shown in the time zone of the conference. The current conference time is: 5th Aug 2026, 05:53:14pm CEST
|
Daily Overview |
| Session | ||
Keynote Talk: Cordelia Schmid (Inria, Google)
Session Topics: Agentic AI, Foundation Models, Generative Models, Graph Neural Networks, Physics-informed Machine Learning, Reinforcement Learning, Probabilistic Methods, Uncertainty Quantification, Audio, Other, Graphs, Image, Multimodal Data, Simulation Data, Tabular Data, Text, Time Series, Video, Other, Core Machine Learning, Aeronautics, Space & Transport, Energy, Earth & Environment, Health, Information, Matter
| ||
| Session Abstract | ||
|
Discover the latest advances in AI for video understanding and vision-language-guided robotics in this keynote presentation. | ||
| Presentations | ||
10:00am - 10:45am
Invited talk ID: 348 / Thu | LAB 10h Keynote_Schmid: 001 Modalities: Multimodal Data, Video Methods: Foundation Models, Physics-informed Machine Learning, Reinforcement Learning Application Domain: Aeronautics, Space & Transport Video-Guided Policies for Robotic Manipulation Inria, France In this talk, we first present a novel approach and benchmark for long-horizon robotic manipulation. Our method integrates the high-level task planning capabilities of Large Language Models (LLMs) with the precise object grounding of Vision-Language Models (VLMs). Given a detailed grounded plan, a 3D low-level motion planner executes actions conditioned on natural language. While our approach demonstrates excellent performance in real-world settings, robust manipulation also requires reliable error handling. We introduce a framework for detecting planning and execution failures, highlighting the critical role of high-quality training data. In challenging long-horizon tasks, this failure detection and recovery mechanism significantly improves overall system performance. In the second part of the talk, we explore learning vision-based policies for multi-fingered robot hands using human video demonstrations. Our method employs reinforcement learning with trajectory-guided rewards and unified visual policy training. Experiments in both simulation and real-world environments demonstrate that our approach outperforms state-of-the-art methods across three complex dexterous manipulation tasks. Because training these models requires high-fidelity visual data, we propose a method to generate temporally dense and consistent captions and object groundings. We conclude by showing that a model trained on a large-scale automatically annotated dataset achieves state-of-the-art results for this task. | ||
