All research
Vision-language navigation
Centre for Systems and Control, IIT Bombay
LiDAR point cloudsvision-language models
Abstract
Navigating from a natural-language instruction means the goal is never given as a coordinate. The robot has to ground phrases against what it perceives, and decide when it has arrived without ever being told where that is.
The problem
The work operates over city-scale LiDAR point clouds, where the geometry is dense and complete but carries no semantics on its own. An instruction refers to landmarks and relations; the map contains only points.