All research

Vision-language navigation

Centre for Systems and Control, IIT Bombay

LiDAR point cloudsvision-language models

Abstract

Navigating from a natural-language instruction means the goal is never given as a coordinate. The robot has to ground phrases against what it perceives, and decide when it has arrived without ever being told where that is.

The problem

The work operates over city-scale LiDAR point clouds, where the geometry is dense and complete but carries no semantics on its own. An instruction refers to landmarks and relations; the map contains only points.