3rd AI Meets Autonomy: Vision, Language, and Autonomous Systems Workshop
IROS 2026, Pittsburgh, Pennsylvania, USA
Thursday, Oct. 1st, 8:30AM - 12:30PM
IROS 2026, Pittsburgh, Pennsylvania, USA
Thursday, Oct. 1st, 8:30AM - 12:30PM
This workshop will be held in the morning of Thursday, Oct. 1, 2026 from 8:30AM - 12:30PM. Location: TBD
Recent advances in Large Language Models (LLMs), Visual Language Models (VLMs), and other foundation models present new opportunities for robotics. The 3rd iteration of this workshop focuses on exploring the intersection of these models and robotic systems, highlighting how progress in the AI and computer vision communities can inform and accelerate robotics research. Integrating LLMs and VLMs into robotic pipelines could enable systems that are more explainable, instructable, and capable of generalizing across tasks. However, achieving seamless integration remains a significant challenge, as existing models often lack the grounded understanding required for real-world robotic applications, including knowledge of physical properties, spatial relationships, and temporal dynamics. Closer integration with robotic platforms may help address these limitations, as real-world interaction provides rich sensory observations and physical feedback that can support the development of more robust and physically grounded intelligent systems.
This workshop seeks to establish an inclusive and collaborative forum for professionals, researchers, and enthusiasts to exchange ideas, share experiences, and build connections within the AI and robotics community, with particular emphasis on supporting and connecting early-career researchers. The program will include invited talks, paper presentations, open panel discussions, networking opportunities, and presentations of exclusive results and demonstrations from the CMU Vision-Language-Autonomy Challenge. Five invited speakers will present their research, perspectives, and future directions on topics at the intersection of AI and autonomous systems, spanning areas such as datasets and benchmarks, software infrastructures, visual-language navigation, situated reasoning, robotics foundation models, and related themes.
See AI Meets Autonomy 2025, AI Meets Autonomy 2024 for the previous iterations of our workshop at IROS 2025 (Hangzhou) and IROS 2024 (Abu Dhabi).
Topics of discussion and open questions include but are not limited to the following:
Vision-Language Navigation
Semantic SLAM
Semantic Mapping
Foundation Models for Robotics
LLMs for Robotics
Human-Robot Interaction
Vision-Language-Action Models
Object-Goal Navigation
Embodied Question Answering
Spatial and Causal Reasoning
Real-Robot Autonomy Stacks
We invite paper submissions to the 3rd AI Meets Autonomy: Vision, Language, and Autonomous Systems Workshop.
This is a non-archival workshop, so both previously published papers and papers currently under review at conferences or other workshops are welcome. Submissions should follow the IROS 2026 paper format and be 4 to 8 pages in length (including references). Submissions do not need to be double-blind; single-blind submissions are sufficient. Submissions may cover a broad range of topics, including (but not limited to) vision-language navigation, semantic SLAM and mapping, foundation models for robotics, LLMs for robotics, human-robot interaction, vision-language-action models, and related areas.
All submissions will undergo peer review. Selected outstanding papers will be invited to give a spotlight presentation at the workshop.
Call Submission Deadline: August 25th, 2026, 23:59 AoE
Submission Portal: OpenReview
Notification of Acceptance: September 5th, 2026
We are now accepting late-breaking submissions to the workshop.
The deadline is September 17, 23:59PM Anywhere on Earth (AoE).
Submissions are single-blind (please include author names in your PDF) and should be submitted via this [Google Form link], not OpenReview.
This is a non-archival track, so you're welcome to reuse or resubmit existing conference submissions (e.g., your ICRA 2027 submission).
No formal reviews will be provided. We will only notify authors of accept (invited to the poster session) or reject decisions. For fairness, late-breaking submissions will not be considered for spotlight presentations.
Please refer to CMU VLN Challenge. The top-performing teams will have the opportunity to present their results at the workshop.
Krishna Murthy Jatavallabhula
Johns Hopkins University
Yonatan Bisk
Carnegie Mellon University
He Wang
Peking University
Bernadette Bucher
University of Michigan
David Fan
FieldAI
CRAFT: Video Diffusion for Bimanual Robot Data Generation. Jason Chen, I-Chun Arthur Liu, Gaurav S. Sukhatme, Daniel Seita
MimicAgent: Learning Quadruped Skills via Text-to-Trajectory Generation. Lucky Kant Nayak, Narayanan Palghat Parameswaran, Neehar Peri, Deva Ramanan
Cross-Environment LiDAR–Camera Misalignment Detection via Object-Level VLM Assessment. Seokhwan Jeong, Younggun Cho
CMU-Drive and V2V-VLA: Cooperative Multi-agent Unified Driving with Reasoning Benchmark and Vehicle-to-Vehicle Vision-Language-Action Models. Hsu-kuang Chiu, Stephen F Smith
Embody Lite: An Auditable Proxy Evaluation of Language-Model Control Interfaces. Jiacheng Liu
Large Reward Models: Generalizable Online Robot Reward Generation with Vision-Language Models. Yanru Wu, Weiduo Yuan, Ang Qi, Vitor Campagnolo Guizilini, Jiageng Mao, Yue Wang
Scalable LLM-based PDDL Domain Generation for Aerial Robotics. Songhao Huang, Yuwei Wu, Guangyao Shi, Gaurav S. Sukhatme, Vijay Kumar.
ODG-NoMaD: Overhead-Camera Direction-Guided NoMaD. Blossom Treesa Bastian, Manish Kolachalam, Rani malhotra, Keerthi S Shetty, Ashish Dutta
A Survey of Agentic Robotics: Toward Continual Self-Improvement. Jiaming Wang, Yuhua Jiang, Liu Diwen, Chen Jizhuo, Zhengcheng Shen, Harold Soh
A Pragmatist Robot: Learning Task Planning by Trial and Error with Memory-Augmented VLMs. Kaixian Qu, Guowei Lan, René Zurbrügg, Changan Chen, Christopher Mower, Haitham Bou Ammar, Marco Hutter
MAPLE: Model-Agnostic LLM Agents for Autonomous Experiments using MADSci. Gabriel Ferrer, William Shields, Turgay Korkmaz
Autonomous Culvert Inspection using VLMs. Yashom Dighe, Yash Turkar, Karthik K Dantu
Evaluating VLMs for Visual Path Planning: A Benchmark and an Overlay Tool. Christine Ohenzuwa, Wenhao Luo, Katia P. Sycara, Woojun Kim
Wenshan Wang
CMU Robotics Institute
Ji Zhang
CMU NREC & Robotics Institute
Avigyan Bhattacharya
CMU Robotics Institute
Seungchan Kim
CMU Robotics Institute
Haochen Zhang
CMU Robotics Institute
Jonas Frey
Stanford University & UC Berkeley
Fadhil Ginting
FieldAI
Chuchu Chen
George Washington University
We would like to thank the following reviewers (listed in alphabetical order) for their valuable time and effort in reviewing the workshop papers.
Bavin Saravanan
Byeonghyun Pak
Dmytro Kurdydyk
Ishaan Malhotra
Jaskaran Singh Sodhi
Jianwen Cao
Junyoung Kim
Krrish Jain
Mangat Rai
Muyang Yan
Pravin Kumar Muralidaran
Raghav Sharma
Sagar Sachdev
Saksham Sharma
Soumya Teotia
Sunghwan Kim
Varun Kandiyappan
Zhizheng Liu
Zuntao Liu