Meanings and Measurements: Multi-Agent Probabilistic Grounding for Vision-Language Navigation
Swagat Padhan, Lakshya Jain, Bhavya Minesh Shah, et al.
This paper addresses the challenge of robots understanding and following human instructions that combine semantic meaning with physical measurements, like 'go two meters to the right of the fridge.' The authors propose MAPG, a system that breaks down complex language instructions into smaller parts, uses AI vision-language models to understand each part, and then combines these interpretations to produce precise, physically grounded robot movements. They demonstrate improvements on benchmark tests and show the approach works on real robots.
vision-language modelsrobot navigationgroundingmulti-agent systems