Title: Compositional Generative Modeling for Robotic Planning and Control

Date: Tuesday, September 1st, 2026

Time: 9:00 AM to 10:30 AM ET

 

In-person Location: CODA C0903 Midtown

Zoom: https://gatech.zoom.us/j/97270969399?pwd=XTTOoLyDVa6w8xOKXcTYLmq6I5YueA.1

 

Utkarsh Aashu Mishra

Ph.D. Student
Institute for Robotics & Intelligent Machines 
Georgia Institute of Technology

 

Committee:

Dr. Yongxin Chen (advisor) -- Daniel Guggenheim School of Aerospace Engineering, Georgia Institute of Technology

Dr. Danfei Xu (advisor) -- School of Interactive Computing, Georgia Institute of Technology

Dr. Harish Ravichandar -- School of Interactive Computing, Georgia Institute of Technology

Dr. Leslie Pack Kaelbling -- Department of Electrical Engineering and Computer Science, Massachusetts Institute of Technology

Dr. Yilun Du -- School of Engineering and Applied Sciences, Harvard University

 

Abstract:

Long-horizon manipulation requires planning for many interdependent actions in sequence to satisfy a goal condition. Current robot-learning methods are limited along three axes: (a) Data: whole-task demonstrations are combinatorially expensive, so monolithic policies do not scale. (b) Inter-step dependencies: a locally optimal early action can make a later step infeasible, and errors compound. (c) Partial observability: an open-world scene is not fully known at planning time. The first part of this proposal develops a long-horizon planning framework that learns distributions of short-horizon behaviors using generative models and composes them at inference to generate long-horizon plans. The compositional generative modeling reasons about inter-step dependencies and goal reaching without any whole-task demonstrations. The second part extends this framework toward open-world planning along two directions: (i) it explores the compositional generalization of short-horizon models trained on a diverse set of real-world, manipulation-rich behaviors, and (ii) it adds a verification harness, an independent symbolic-geometric verifier that checks whether a proposed plan is geometrically executable before execution. Across both axes, the objective is to build a reliable framework that can replan and sufficiently generalize across behaviors and planning horizon: a composed plan should be a plausible real-world future prediction while satisfying the long-horizon goal under inter-step constraints.