Zero-shot Object-Centric Instruction Following: Integrating Foundation Models with Traditional Navigation
LIFGIF grounds object-centric language instructions in a factor-graph map and was evaluated in simulation and on a Boston Dynamics Spot.
Why it matters
Large scale scenes such as multifloor homes can be robustly and efficiently mapped with a 3D graph of landmarks estimated jointly with robot poses in a factor graph, a technique commonly used in commercial robots such as drones and robot vacuums. In this work, we propose Language-Inferred Factor Graph for Instruction Following (LIFGIF), a zero-shot method to ground natural language instructions in such a map.

Evidence notes
- The work releases OC-VLN, a dataset of landmark-referenced navigation instructions, and compares LIFGIF with two zero-shot baselines.
- LIFGIF outperformed both baselines across Success Rate, SPL, Oracle Success Rate and normalized Dynamic-Time Warping.
- The real-world demonstration used Spot's SDK for waypoint navigation, RTAB-Map for localization and YOLOv7 in an office environment.
Company context
Boston Dynamics builds high-performance mobile robots including Spot for industrial and field inspection, Stretch for warehouse trailer unloading, and Atlas for advanced humanoid mobility.